fix(ai-red-teaming): record OWASP-Agentic slug as airt_goal_category - #161
Merged
Merged
Conversation
Agentic attacks resolved their goal_category through GOAL_CATEGORY_ALIASES to an OWASP-LLM GoalCategory enum and recorded that enum value, so findings mapped to LLM01/jailbreak_general instead of the correct ASI category. Keep the canonical agentic slug verbatim in a new AIRT_GOAL_CATEGORY constant (falling back to GOAL_CATEGORY.value for non-agentic attacks) so the platform maps them to OWASP-Agentic (ASI). Also resolve agentic slugs to a valid SDK enum for scorers instead of falling back to JAILBREAK_GENERAL.
1 task
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Agentic attacks generated by
attack_runner.py(SDKgenerate_attack/generate_agentic_attackand the TUI agent) were mis-categorized in the platform: findings mapped to OWASP-LLMjailbreak_general(LLM01) or a collapsed LLM enum instead of the correct OWASP-Agentic (ASI) category.Root cause
goal_categorywas resolved throughGOAL_CATEGORY_ALIASESto an SDKGoalCategoryenum (e.g.agentic_memory_poison -> HARMFUL_CONTENT), and the generated workflow recordedairt_goal_category=GOAL_CATEGORY.value. The agentic identity was lost before it reached ClickHouse, whereapp/airt/clickhouse.py::_OWASP_AGENTICmaps the canonical agentic slug to an ASI category. Additionally, the canonical long-form slugs (agentic_memory_poisoning) weren't in the alias table, so they fell back toJAILBREAK_GENERALwith a warning.Fix
AGENTIC_GOAL_CATEGORIES(canonical platform slugs, keys aligned with_OWASP_AGENTIC) +AGENTIC_GOAL_CATEGORY_ALIASES(short/legacy -> canonical)._resolve_airt_goal_category()returns the canonical agentic slug (orNonefor non-agentic).AIRT_GOAL_CATEGORYconstant: the agentic slug verbatim when agentic, elseGOAL_CATEGORY.value. Emission sites reference the constant._resolve_goal_category()now maps agentic slugs to a valid SDK enum (for scorers) instead ofJAILBREAK_GENERAL.GOAL_CATEGORY.value).Net effect:
agentic_memory_poisoning -> ASI06,agentic_tool_misuse -> ASI02,agentic_supply_chain -> ASI10, etc., in findings/compliance mapping.Validation
pre-commit run --files ...(ruff, ruff-format, check-yaml, gitleaks): pass.just validate->ai-red-teaming@1.17.4OK.1.17.3 -> 1.17.4.Test plan
1.17.4; run an agentic attack via TUI/SDK and confirm the finding shows the ASI category (e.g.agentic_memory_poisoning-> ASI06) instead ofjailbreak_general.