Skip to content

fix(ai-red-teaming): record OWASP-Agentic slug as airt_goal_category - #161

Merged
rdheekonda merged 1 commit into
mainfrom
fix/agentic-goal-category-recording
Sep 30, 2026
Merged

rdheekonda merged 1 commit into
mainfrom
fix/agentic-goal-category-recording

Conversation

@rdheekonda

Copy link
Copy Markdown
Contributor

Summary

Agentic attacks generated by attack_runner.py (SDK generate_attack/generate_agentic_attack and the TUI agent) were mis-categorized in the platform: findings mapped to OWASP-LLM jailbreak_general (LLM01) or a collapsed LLM enum instead of the correct OWASP-Agentic (ASI) category.

Root cause

goal_category was resolved through GOAL_CATEGORY_ALIASES to an SDK GoalCategory enum (e.g. agentic_memory_poison -> HARMFUL_CONTENT), and the generated workflow recorded airt_goal_category=GOAL_CATEGORY.value. The agentic identity was lost before it reached ClickHouse, where app/airt/clickhouse.py::_OWASP_AGENTIC maps the canonical agentic slug to an ASI category. Additionally, the canonical long-form slugs (agentic_memory_poisoning) weren't in the alias table, so they fell back to JAILBREAK_GENERAL with a warning.

Fix

  • Add AGENTIC_GOAL_CATEGORIES (canonical platform slugs, keys aligned with _OWASP_AGENTIC) + AGENTIC_GOAL_CATEGORY_ALIASES (short/legacy -> canonical).
  • New _resolve_airt_goal_category() returns the canonical agentic slug (or None for non-agentic).
  • Generated workflows now emit an AIRT_GOAL_CATEGORY constant: the agentic slug verbatim when agentic, else GOAL_CATEGORY.value. Emission sites reference the constant.
  • _resolve_goal_category() now maps agentic slugs to a valid SDK enum (for scorers) instead of JAILBREAK_GENERAL.
  • Non-agentic attacks are unchanged (fallback to GOAL_CATEGORY.value).

Net effect: agentic_memory_poisoning -> ASI06, agentic_tool_misuse -> ASI02, agentic_supply_chain -> ASI10, etc., in findings/compliance mapping.

Validation

  • Offline: resolvers + generated CONFIG section + attack-param emission asserted (agentic slug recorded verbatim; non-agentic falls back; scorer enum valid).
  • pre-commit run --files ... (ruff, ruff-format, check-yaml, gitleaks): pass.
  • just validate -> ai-red-teaming@1.17.4 OK.
  • Version bumped 1.17.3 -> 1.17.4.

Test plan

  • Publish 1.17.4; run an agentic attack via TUI/SDK and confirm the finding shows the ASI category (e.g. agentic_memory_poisoning -> ASI06) instead of jailbreak_general.

Agentic attacks resolved their goal_category through GOAL_CATEGORY_ALIASES to
an OWASP-LLM GoalCategory enum and recorded that enum value, so findings mapped
to LLM01/jailbreak_general instead of the correct ASI category. Keep the
canonical agentic slug verbatim in a new AIRT_GOAL_CATEGORY constant (falling
back to GOAL_CATEGORY.value for non-agentic attacks) so the platform maps them
to OWASP-Agentic (ASI). Also resolve agentic slugs to a valid SDK enum for
scorers instead of falling back to JAILBREAK_GENERAL.
@rdheekonda
rdheekonda merged commit 9a1a685 into main Sep 30, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant