FEAT: attempt correlation across an attack's conversations (phase 9) - #2946
Conversation
|
Can we put Each conversation should belong to one attack execution. An execution can own several conversations—objective, adversarial, scoring, converter, and branches—but a new execution that uses existing history should copy it into a new conversation, not reuse the original. Messages already link to their conversation through The early allocation of the result ID makes sense; I'd keep that. Please make the conversation-level relationship explicit and queryable, and enforce that an existing conversation cannot be assigned to a different execution. Copies made within an execution should retain that execution's ID; copies used by a new execution should receive the new ID. Conversations created by child attacks should link to the child execution. Please preserve the ability for targets and harnesses to access the ID before sending, including scorer, converter, and streaming calls. That does not require persisting it on every message. Tracking whether a particular message was actually sent, rather than copied as history, is a separate provenance concern. |
An attack now allocates the ID of its AttackResult when execution starts. The prompt normalizer records that ID in the metadata of every request persisted while the attack runs, across the main, adversarial and scoring conversations, and the completed or error result is stored under the same ID. Copied prepended history drops the key, so it never names another attack's result. Targets see the ID on each request, which gives a harness a value to stamp onto the evidence it emits.
The AttackResult ID is still allocated when execution starts, but the link now lives on Conversation and a new indexed attack_result_id column on the Conversations table, added by an Alembic migration. Conversations registered during an execution are linked to it, and assigning a conversation to a different execution raises ValueError. Copies made within an execution keep its ID, history reused by a new execution is copied into a new conversation, and child attacks link their own conversations. Targets, scorers and converters read the current ID through get_current_attack_result_id(). The per-piece prompt_metadata stamping is removed.
7e0e2ca to
3adb4b5
Compare
|
Richard Lundeen (@richlundeen) Thanks for the review, that makes the model much cleaner. The ID is still allocated when execution starts, but the link now lives on Conversation as attack_result_id, stored in a new indexed column on the Conversations table with an Alembic migration, and an existing conversation that is assigned to a different execution raises a ValueError (re-registering with the same ID is a no-op). Copies made within an execution keep its ID, history used by a new execution is copied into a new conversation that takes the new ID, and conversations created by a child attack such as those in SequentialAttack link to the child's result. Targets, scorers, converters and streaming code read the current ID before sending through get_current_attack_result_id() in pyrit.models, which is backed by the existing context variable. The per-message prompt_metadata stamping is removed, and memory can now return an execution's conversations or pieces directly by result ID. |
Link manual attacks and branches to their execution, claim unowned conversations atomically, and move execution scope into common infrastructure. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Description
Implements phase 9 of the scorer contract proposal.
An attack allocates the ID of its
AttackResultonce, when execution starts, and links every conversation it creates or uses to that execution.attack_result_idis already a client-generated UUID and the primary key ofAttackResultEntry, so preallocating it gives each execution one identity, and the link is the row a scorer or harness loads.Conversation.attack_result_id, stored in a new indexed, nullable column on theConversationstable with an Alembic migration.ValueError.SequentialAttack, link to the child's result.get_current_attack_result_id()inpyrit.models, backed by a context variable, so concurrent executions stay separate. Nothing is stamped on individual message pieces.get_attack_result_conversations_async, andget_message_pieces_asyncaccepts anattack_result_idfilter.Scoring behaviour is unchanged.
Tests and Documentation
New tests in
tests/unit/executor/attack/core/test_attack_result_correlation.pyandtests/unit/memory/memory_interface/test_interface_conversation_attack_link.pycover: one ID per execution, persisted and distinct across executions; objective, adversarial (RedTeamingAttack), scorer and branch-copy conversations linked; reassignment to another execution raising; prepended history copied into a new conversation with the new ID; aSequentialAttackchild linking to the child's result; a target and a converter reading the ID before sending; the error result keeping the allocated ID; concurrent executions keeping their own IDs.tests/unit/memory/test_migration.pyupgrades and downgrades the new migration on SQLite.Each behaviour fails without its change: removing the link on registration, the reassignment check, the execution scope, or the nested child scope each fails the tests aimed at it.
python -m pytest tests/unit/executor tests/unit/prompt_normalizer tests/unit/models tests/unit/score -n 8
6189 passed, 78 skipped
python -m pytest tests/unit/memory tests/unit/executor/attack/core/test_attack_result_correlation.py -n 8
1052 passed, 1 skipped
ruff check,ruff format --checkandty checkare clean on the touched files. The branch is rebased onto current main and the migration follows the current head. Documented indoc/code/memory/3_memory_data_types.md.