Skip to content

FEAT: attempt correlation across an attack's conversations (phase 9) - #2946

Merged
Richard Lundeen (richlundeen) merged 3 commits into
microsoft:mainfrom
WatchTree-19:feat/attack-attempt-correlation
Oct 1, 2026
Merged

Richard Lundeen (richlundeen) merged 3 commits into
microsoft:mainfrom
WatchTree-19:feat/attack-attempt-correlation

Conversation

@WatchTree-19

@WatchTree-19 WatchTree-19 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Description

Implements phase 9 of the scorer contract proposal.

An attack allocates the ID of its AttackResult once, when execution starts, and links every conversation it creates or uses to that execution. attack_result_id is already a client-generated UUID and the primary key of AttackResultEntry, so preallocating it gives each execution one identity, and the link is the row a scorer or harness loads.

  • The link lives on Conversation.attack_result_id, stored in a new indexed, nullable column on the Conversations table with an Alembic migration.
  • A conversation registered during an execution takes that execution's ID. Re-registering with the same ID is a no-op, and assigning an existing conversation to a different execution raises ValueError.
  • Copies made within an execution keep its ID. History used by a new execution is copied into a new conversation that takes the new ID. Conversations created by a child attack, such as those in SequentialAttack, link to the child's result.
  • Targets, scorers, converters and streaming code can read the current ID before sending through get_current_attack_result_id() in pyrit.models, backed by a context variable, so concurrent executions stay separate. Nothing is stamped on individual message pieces.
  • Memory can return an execution's conversations with get_attack_result_conversations_async, and get_message_pieces_async accepts an attack_result_id filter.

Scoring behaviour is unchanged.

Tests and Documentation

New tests in tests/unit/executor/attack/core/test_attack_result_correlation.py and tests/unit/memory/memory_interface/test_interface_conversation_attack_link.py cover: one ID per execution, persisted and distinct across executions; objective, adversarial (RedTeamingAttack), scorer and branch-copy conversations linked; reassignment to another execution raising; prepended history copied into a new conversation with the new ID; a SequentialAttack child linking to the child's result; a target and a converter reading the ID before sending; the error result keeping the allocated ID; concurrent executions keeping their own IDs. tests/unit/memory/test_migration.py upgrades and downgrades the new migration on SQLite.

Each behaviour fails without its change: removing the link on registration, the reassignment check, the execution scope, or the nested child scope each fails the tests aimed at it.

python -m pytest tests/unit/executor tests/unit/prompt_normalizer tests/unit/models tests/unit/score -n 8
6189 passed, 78 skipped

python -m pytest tests/unit/memory tests/unit/executor/attack/core/test_attack_result_correlation.py -n 8
1052 passed, 1 skipped

ruff check, ruff format --check and ty check are clean on the touched files. The branch is rebased onto current main and the migration follows the current head. Documented in doc/code/memory/3_memory_data_types.md.

@richlundeen

Copy link
Copy Markdown
Contributor

Can we put attack_result_id on Conversation rather than in each message piece's prompt_metadata? This is the main change I'd like before merging.

Each conversation should belong to one attack execution. An execution can own several conversations—objective, adversarial, scoring, converter, and branches—but a new execution that uses existing history should copy it into a new conversation, not reuse the original. Messages already link to their conversation through conversation_id, so the execution relationship belongs on the existing Conversation model and Conversations table.

The early allocation of the result ID makes sense; I'd keep that. Please make the conversation-level relationship explicit and queryable, and enforce that an existing conversation cannot be assigned to a different execution. Copies made within an execution should retain that execution's ID; copies used by a new execution should receive the new ID. Conversations created by child attacks should link to the child execution.

Please preserve the ability for targets and harnesses to access the ID before sending, including scorer, converter, and streaming calls. That does not require persisting it on every message. Tracking whether a particular message was actually sent, rather than copied as history, is a separate provenance concern.

An attack now allocates the ID of its AttackResult when execution starts.
The prompt normalizer records that ID in the metadata of every request
persisted while the attack runs, across the main, adversarial and scoring
conversations, and the completed or error result is stored under the same
ID. Copied prepended history drops the key, so it never names another
attack's result. Targets see the ID on each request, which gives a harness
a value to stamp onto the evidence it emits.
The AttackResult ID is still allocated when execution starts, but the link
now lives on Conversation and a new indexed attack_result_id column on the
Conversations table, added by an Alembic migration. Conversations registered
during an execution are linked to it, and assigning a conversation to a
different execution raises ValueError. Copies made within an execution keep
its ID, history reused by a new execution is copied into a new conversation,
and child attacks link their own conversations. Targets, scorers and
converters read the current ID through get_current_attack_result_id(). The
per-piece prompt_metadata stamping is removed.
@WatchTree-19
WatchTree-19 force-pushed the feat/attack-attempt-correlation branch from 7e0e2ca to 3adb4b5 Compare October 1, 2026 20:55
@WatchTree-19

Copy link
Copy Markdown
Contributor Author

Richard Lundeen (@richlundeen) Thanks for the review, that makes the model much cleaner. The ID is still allocated when execution starts, but the link now lives on Conversation as attack_result_id, stored in a new indexed column on the Conversations table with an Alembic migration, and an existing conversation that is assigned to a different execution raises a ValueError (re-registering with the same ID is a no-op). Copies made within an execution keep its ID, history used by a new execution is copied into a new conversation that takes the new ID, and conversations created by a child attack such as those in SequentialAttack link to the child's result. Targets, scorers, converters and streaming code read the current ID before sending through get_current_attack_result_id() in pyrit.models, which is backed by the existing context variable. The per-message prompt_metadata stamping is removed, and memory can now return an execution's conversations or pieces directly by result ID.

Link manual attacks and branches to their execution, claim unowned conversations atomically, and move execution scope into common infrastructure.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@richlundeen
Richard Lundeen (richlundeen) added this pull request to the merge queue Oct 1, 2026
Merged via the queue into microsoft:main with commit 016d9b2 Oct 1, 2026
52 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants