Repository navigation
FIX: degrade an unreadable Responses output to an empty marker - #2906
hannahwestra25 merged 12 commits into
Conversation
_construct_message_from_response_async tracks has_visible_response but only consulted it on the truncated path, so a completed response whose output PyRIT cannot read came back as a successful Message holding the reasoning dump. With a built-in tool enabled (image_generation, code_interpreter, file_search, ...) the Responses API returns sections such as image_generation_call, which _parse_response_output_section skips with `return None`; the run then scored the reasoning JSON as the model's answer with response_error="none". OpenAIChatTarget already raises EmptyResponseException when a response that is not truncated yields no content. Do the same here, and keep returning the message when a readable section is present next to an unmodelled one.
`_send_model_request_async` is wrapped in `@pyrit_target_retry`, which retries `RateLimitError | EmptyResponseException | RateLimitException`. The check added in this branch raised `EmptyResponseException` for a response that *completed* with no section PyRIT models, so every retry reproduced the same shape: ten billed calls plus backoff before the agentic loop gave up, and the tool messages it had collected were dropped with the exception. `doc/contributing/9_exception.md` scopes retry to rate limits and parse failures, so raise `PyritException` instead. The docstring now says so rather than naming an exception that is no longer raised. The test asserts `type(excinfo.value) is PyritException` and that it is not an `EmptyResponseException`. Asserting only `PyritException` would not have caught this, since `EmptyResponseException` subclasses it -- the weaker assertion passes on the old code too. Reported by @hannahwestra25.
|
Right, and I checked the mechanism before changing it: Done in One thing worth flagging because my first attempt got it wrong: asserting only with pytest.raises(PyritException) as excinfo:
...
assert not isinstance(excinfo.value, EmptyResponseException)
assert type(excinfo.value) is PyritExceptionwhich fails on One asymmetry I did not change, because it is outside this PR and may be deliberate: |
Raising made the batch fail and skipped metadata capture, and the proposed PyritException was retried by pyrit_target_retry on a deterministic outcome. Append the graceful empty marker instead, keep reasoning pieces, and pin the shape with a reasoning-only regression test.
|
Agreed — went with the append-marker shape instead of raising. The completed-but-unreadable path now appends |
|
Thanks hannahwestra25, both applied:
|
The completed path warns and the truncated path stays quiet, but neither outcome was asserted, so dropping the warning or the truncation gate would both regress silently. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Description
OpenAIResponseTargetreports success for a completed Responses call whose output carries nothing PyRIT can read. With a built-in tool enabled (image_generation,code_interpreter,file_search, …) the API answers with sections such asimage_generation_call, which_parse_response_output_sectionskips. The target then returns aMessagewhose only piece is the reasoning dump, withresponse_error="none", so the attack loop records that JSON as the model's answer.has_visible_responseis already computed while looping over the sections, but was only consulted on the truncated path. It is now consulted on every path: when nothing visible was produced, a graceful empty marker piece (response_error="empty") is appended viabuild_empty_truncated_response, matching whatOpenAIChatTargetandLiteLLMChatTargetalready do on truncation. Reasoning pieces are retained after it for memory and debugging, and a readable section sitting next to an unmodelled one is still returned unchanged.Degrading rather than raising is deliberate.
EmptyResponseExceptionis inpyrit_target_retry's retry set, so raising would re-send a request whose outcome is deterministic, and an exception would also skip the_capture_response_metadatacall below it, losing token counts for a response the provider already billed. A genuinely empty output still raisesEmptyResponseExceptionfrom_response_adapter.validate(), which runs before this method.The completed path also logs a warning, since the empty marker is otherwise the operator's only signal that a run is producing nothing readable. The truncated path stays quiet: hitting the token cap is an expected outcome and was already silent.
Tests and Documentation
Five cases in
tests/unit/prompt_target/target/test_openai_response_target.py:mainmainmainThe
_construct_message_from_response_asyncdocstring no longer scopes the empty-piece fallback to truncation.pytest tests/unit/prompt_target/target/test_openai_response_target.py -qgives 114 passed.ruff checkandruff format --checkare clean on both files, andpre-commit run --filespasses on the test file.