FIX: raise when a completed Responses output has nothing readable - #2906
fei (feiiiiii5) wants to merge 3 commits into
Conversation
_construct_message_from_response_async tracks has_visible_response but only consulted it on the truncated path, so a completed response whose output PyRIT cannot read came back as a successful Message holding the reasoning dump. With a built-in tool enabled (image_generation, code_interpreter, file_search, ...) the Responses API returns sections such as image_generation_call, which _parse_response_output_section skips with `return None`; the run then scored the reasoning JSON as the model's answer with response_error="none". OpenAIChatTarget already raises EmptyResponseException when a response that is not truncated yields no content. Do the same here, and keep returning the message when a readable section is present next to an unmodelled one.
| # A response that completed without a readable section is a failure, not an answer. | ||
| # The chat target raises in the same situation; reporting reasoning or a section | ||
| # type PyRIT does not model as the model's response would let it be scored as one. | ||
| raise EmptyResponseException(message="Failed to extract any response content.") |
There was a problem hiding this comment.
EmptyResponseException is retried by @pyrit_target_retry, but an unmodelled section type recurs every attempt — so this costs 10 billed calls plus backoff, and drops the agentic loop's collected tool messages. Per doc/contributing/9_exception.md, deterministic failures shouldn't retry.
| raise EmptyResponseException(message="Failed to extract any response content.") | |
| raise PyritException(message="Failed to extract any readable response content.") |
`_send_model_request_async` is wrapped in `@pyrit_target_retry`, which retries `RateLimitError | EmptyResponseException | RateLimitException`. The check added in this branch raised `EmptyResponseException` for a response that *completed* with no section PyRIT models, so every retry reproduced the same shape: ten billed calls plus backoff before the agentic loop gave up, and the tool messages it had collected were dropped with the exception. `doc/contributing/9_exception.md` scopes retry to rate limits and parse failures, so raise `PyritException` instead. The docstring now says so rather than naming an exception that is no longer raised. The test asserts `type(excinfo.value) is PyritException` and that it is not an `EmptyResponseException`. Asserting only `PyritException` would not have caught this, since `EmptyResponseException` subclasses it -- the weaker assertion passes on the old code too. Reported by @hannahwestra25.
|
Right, and I checked the mechanism before changing it: Done in One thing worth flagging because my first attempt got it wrong: asserting only with pytest.raises(PyritException) as excinfo:
...
assert not isinstance(excinfo.value, EmptyResponseException)
assert type(excinfo.value) is PyritExceptionwhich fails on One asymmetry I did not change, because it is outside this PR and may be deliberate: |
Description
OpenAIResponseTargetreports success for a completed Responses call whose output carries nothing PyRIT can read. With a built-in tool enabled (image_generation,code_interpreter,file_search, …) the API answers with sections such asimage_generation_call, which_parse_response_output_sectionskips. The target then returns aMessagewhose only piece is the reasoning dump, withresponse_error="none", so the attack loop records that JSON as the model's answer and scores it.has_visible_responseis already computed while looping over the sections, but was only consulted on the truncated path.OpenAIChatTarget._construct_message_from_response_asyncraisesEmptyResponseExceptionin this same situation; this makes the Responses target match it, while still returning the message when a readable section sits next to an unmodelled one. The method docstring now documents the raise.Tests and Documentation
Two cases in
tests/unit/prompt_target/target/test_openai_response_target.py: a completed response of reasoning plus an unmodelled section now raisesEmptyResponseException(fails onmainwithDID NOT RAISE), and a readable message next to an unmodelled section is still returned.pytest tests/unit/prompt_target/ -qgives 1494 passed;ruff checkandruff format --checkare clean on both files.