Repository navigation
FEAT Add explicit Scenario progress result roles - #2997
Open
Dev3 ^~^ (shashank03-dev) wants to merge 5 commits into
Open
Dev3 ^~^ (shashank03-dev) wants to merge 5 commits into
Dev3 ^~^ (shashank03-dev) wants to merge 5 commits into
Conversation
Scenario progress now says what each persisted result represents, so consumers no longer interpret class names, technique names, or empty conversations. - AttackResultRole: every AttackStrategy declares RESULT_ROLE (target_facing by default, orchestration for SequentialAttack). It is copied onto the context and written to attribution_data["result_role"] for completed and error results. No migration; rows without it read as unknown, and nothing is inferred from an empty conversation ID. - ScenarioRunPlanGroupKind: AtomicAttack takes a group_kind (attack, baseline from build_baseline_atomic_attack, adaptive from AdaptiveScenario), recorded in the run plan. It is not part of any identity, so runs resume unchanged. Plans saved before this change round-trip without a kind and read as unknown. - SequentialAttack gives each child its 1-based attempt_index, matching the existing _adaptive_attempt label. - The progress query projects attack_metadata, and ScenarioProgressResult exposes result_role, the stored ordered child_attack_result_ids, and attempt_index. ScenarioAtomicGroupProgress exposes kind. Planned-unit completion, retry and error counters, success rates, and preparation-failure semantics are unchanged. Fixes microsoft#2992
Roman Lutz (romanlutz)
requested changes
Oct 6, 2026
_apply_attribution returned early when the context had no Scenario attribution, so standalone execute_async calls stored results with no result_role and read back as unknown. Always stamp result_role, and only set attribution_parent_id and the parent linkage fields when attribution is present. Add real-SQLite tests for standalone PromptSendingAttack and SequentialAttack on both the completed and error paths.
…t-roles # Conflicts: # pyrit/backend/services/scenario_progress_read_model.py
attempt_index matches the _adaptive_attempt label only for techniques that Adaptive runs directly. When a selected technique is itself a compound attack, its children are numbered under that nested parent while their label still names the outer Adaptive attempt. Qualify the statement in both paired adaptive scenario docs and add a regression that runs a nested SequentialAttack through AdaptiveTechniqueDispatcher, a real Scenario and SQLite, and pins the parent-relative index through the progress projection.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixes #2992 (milestone 7 of the Scenario GUI follow-up roadmap).
Scenario progress now says what each persisted result represents, without consumers interpreting class names, technique names, or empty conversations. The producers record the contract and the existing progress projection exposes it in the same change. Counts are untouched.
Result roles, recorded by the producer.
AttackResultRole(pyrit.models) istarget_facing,orchestration, orunknown. EveryAttackStrategydeclaresRESULT_ROLEas a class constant, following the existingDELEGATES_SCORINGpattern: the default isTARGET_FACING, andSequentialAttackdeclaresORCHESTRATION.execute_with_context_asynccopies it onto the context, and_apply_attributionwrites it toattribution_data["result_role"]for both completed and error results. It lives in the existing JSON column, so there is no migration, and milestone 8 can filter on it in SQL.Planned group kinds.
AtomicAttacktakes agroup_kind(ScenarioRunPlanGroupKind):attackby default,baselinefrombuild_baseline_atomic_attack(all 12 baseline call sites go through it), andadaptivefromAdaptiveScenario. The run plan records it asScenarioRunPlanAtomicGroup.kind. It describes the group only: it is not part of the atomic group ID, the technique eval hash, or the scenario identity, so existing runs resume unchanged. The plan field is optional, so plans saved before this change round-trip without gaining a kind (the scenario history projection returns them unchanged), and progress reports their groups asunknown.Ordered children. Orchestration rows expose the
metadata["child_attack_result_ids"]thatSequentialAttackalready persists, in stored order. The progress query now projectsattack_metadata; there is no second hierarchy store.Attempt indexes.
SequentialAttackgives each child a copy of its own attribution with a 1-basedattempt_index, persisted on the child's result row. For Adaptive this equals the existing_adaptive_attemptlabel, which is unchanged. I implemented the genericSequentialAttackposition I proposed in the issue, 1-based to match that label; it is a one-line change if you prefer otherwise.Progress API.
ScenarioProgressResultgainsresult_role,child_attack_result_ids, andattempt_index;ScenarioAtomicGroupProgressgainskind. All four are consumed in this change by the read model.Legacy records. A row without a recognized role reads as
unknown, and nothing is inferred from an empty conversation ID. That heuristic is also wrong for new data: when aSequentialAttackfails, its error result gets a generated conversation ID (covered by a test). Plans saved before this change read their groups asunknown. Malformed child IDs or attempt indexes read as empty orNone.Unchanged: planned-unit completion, retry and error counters, success rates, target-attempt accounting (milestone 8), and preparation-failure semantics. A parent and its children still map to one planned unit.
Out of scope: history aggregates (milestone 8), hierarchy endpoints (milestone 9), and frontend rendering (milestone 10).
frontend/src/types/index.tsis unchanged because no frontend code reads these fields yet. This does not touch the read interfaces that #2953 changes; the read-model edits are confined to_map_progress_delta, three static readers, and one keyword in the atomic-group summary.Proof
The same real run on
main(e7d2619) and on this branch, read throughScenarioProgressReadModel: a baseline and an ordinary attack over two objectives, plus two Adaptive-style objectives, each aSequentialAttackparent with twoPromptSendingAttackchildren. It usesMockPromptTargetand a scorer that never matches, so every attempt fails and both children run. No model calls.mainorchestration, 2 ordered childrentarget_facing,attempt_index1 or 2target_facingEvery count in the progress summary is identical: overall
planned=6 completed=6 errors=0 retries=4 succeeded=0, the same per group, technique, display group, and seed group, andunattributed_attempts=0.This also shows what milestone 8 can now fix. Each Adaptive objective reports
retries=2, onmainand here alike, because the parent and both children map to one planned unit and the extra rows count as re-attempts. Withresult_role, milestone 8 can count target attempts without that inflation. This PR leaves it unchanged, as the issue requires.Tests and Documentation
Tests use real domain objects and the real SQLite memory, with no live target calls:
tests/unit/backend/test_scenario_progress_read_model.py): a realScenariobuilds and persists its plan, runs a baseline, an ordinary attack, and an Adaptive-style group (aSequentialAttackwith twoPromptSendingAttackchildren) againstMockPromptTarget, and the read model projects the stored rows. It asserts the group kinds in the persisted plan and the progress summary; the parent's role, empty conversation, and children in dispatch order after the database round trip; the children's roles and attempt indexes 1 and 2; that_adaptive_attempton each stored child equals its index; that the parent and children still share one planned unit; and the JSON wire values.SequentialAttackerror result is stillorchestrationalthough it has a conversation ID.unknownkind.AttackPreparationFailurekeeps both signals intact and its outcome.SequentialAttackchild gets its 1-based position while the parent's own attribution is untouched;TextAdaptive's plan recordsbaselineandadaptivekinds.result_rolein the stamped dict, the attribution-forwarding test expects the child's copy withattempt_index=1, and theMagicMock(spec=AtomicAttack)fixtures in three core scenario test files setgroup_kind, as every realAtomicAttackhas one.Each guarantee is pinned: deliberately removing
SequentialAttack.RESULT_ROLE, dropping the child position, dropping the group kind from the read model, or inferringorchestrationfrom a missing role each fails the matching tests.doc/code/scenarios/3_adaptive_scenarios.pyand.ipynbdocument the role contract with an example progress response.Checks run locally:
uv run pre-commit run --files <changed files>: all hooks passuv run ty check pyrit tests/unit: no new diagnostics; the output matchesmainapart from line numberspytest -n 2 tests/unit/scenario tests/unit/backend tests/unit/executor tests/unit/memory tests/unit/models: 7539 passed, 39 skippedmain: 100% (52 of 52 changed lines)