Skip to content

fix(ai-red-teaming): record custom-target attacks via assessment - #162

Merged
rdheekonda merged 1 commit into
mainfrom
fix/agentic-custom-target-recording
Sep 30, 2026
Merged

rdheekonda merged 1 commit into
mainfrom
fix/agentic-custom-target-recording

Conversation

@rdheekonda

Copy link
Copy Markdown
Contributor

Summary

Capability wiring for bug #1 (custom-target/multistep attacks recorded 0 spans / 0 findings). multistep_tool_attack, agentvigil_attack, and eva_attack ran their custom loops and returned a report without going through the assessment, so nothing was traced or materialized even on success.

Each generated workflow now calls the new SDK method assessment.record_attack_result(...) right after the attack, so the outcome emits a study + trial span linked to the assessment and materializes a finding with the right OWASP-Agentic category:

  • multistep_tool_attack -> agentic_data_exfil
  • agentvigil_attack / eva_attack -> agentic_goal_hijacking

Dependency / ordering

Requires the SDK method from dreadnode-tiger#2668 at runtime. That SDK PR must merge + release before this capability version is published, or the generated scripts will AttributeError on record_attack_result.

Stacked on fix/agentic-goal-category-recording (PR #161) to avoid a capability.yaml version-bump conflict; base retargets to main once #161 merges.

Validation

  • All three generators produce compilable scripts (internal compile() passes with the new calls).
  • pre-commit (ruff, ruff-format, yaml, gitleaks): pass.
  • just validate -> ai-red-teaming@1.17.5 OK.
  • Version 1.17.4 -> 1.17.5.

Test plan

  • After SDK #2668 releases, run each attack against a prod agent target; confirm a study:/trial span in ClickHouse for the assessment and a materialized finding in the UI (correct category + score). Honeytoken flows (TUI tools) are out of scope here - tracked separately.

@rdheekonda
rdheekonda changed the base branch from fix/agentic-goal-category-recording to main September 30, 2026 00:48
multistep_tool_attack, agentvigil_attack and eva_attack ran their custom loops
and returned a report without going through the assessment, so they emitted no
study/trial spans and materialized no findings even on success. Call the new
assessment.record_attack_result(...) after each so the outcome is traced and
mapped to the right OWASP-Agentic category (data exfil / goal hijacking).

Requires the SDK method from dreadnode-tiger#2668.
@rdheekonda
rdheekonda force-pushed the fix/agentic-custom-target-recording branch from dbdb6b7 to 1c2b800 Compare September 30, 2026 00:53
@rdheekonda
rdheekonda merged commit be7341b into main Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant