Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion capabilities/ai-red-teaming/agents/ai-red-teaming-agent.md
Original file line number Diff line number Diff line change
Expand Up @@ -157,7 +157,7 @@ The AI Red Teaming capability provides these tools:
- **generate_attack** — Generate + auto-execute an attack workflow (single, campaign, or transform study)
- **generate_category_attack** — Generate + auto-execute a category-based assessment from bundled goals
- **generate_agentic_attack** — Generate + auto-execute a single attack against an HTTP agent API
- **generate_multistep_tool_attack** — Generate + auto-execute an algorithmic multi-step tool-attack search against an HTTP agent: finds replay-stable read-then-exfiltrate causal paths (agent reads an untrusted source, then a privileged outbound action carries its contents). Adapts the Kaggle AI Agent Security - Multi-Step Tool Attacks techniques (replay-stable candidates, shared-prefix branching, two-probe recovery, split/multi-post exfil). Pass `agent_url` (+ optional `secret_source` / `sink_url` / `sink_email` / `time_budget_s`)
- **generate_multistep_tool_attack** — Generate + auto-execute an algorithmic multi-step tool-attack search against an HTTP agent: finds replay-stable read-then-exfiltrate causal paths (agent reads an untrusted source, then a privileged outbound action carries its contents). Black-box techniques: replay-stable candidates, shared-prefix branching, calibrated two-probe recovery, split/multi-post exfil. Pass `agent_url` (+ optional `secret_source` / `sink_url` / `sink_email` / `time_budget_s`)
- **generate_agentic_suite_attack** — Generate + auto-execute the FULL agentic suite against an HTTP agent API: every OWASP-ASI category, auto-selecting the mapped attacks + family transforms (MCP, multi-agent, reasoning, exfiltration, …) + detection scorers. The "run all possible attacks on my agent" path — no need to name individual attacks; omit `categories` for everything.
- **generate_agentvigil_attack** — Generate + auto-execute an AgentVigil MCTS indirect-injection search against an HTTP agent: plants injection payloads in the agent's input across a set of benign tasks and searches (UCB1 selection, reward = ASR + coverage, five LLM mutation operators) for a payload that reliably fires the `instructed_tool`. Evidence-gated on the tool actually firing. Our implementation of arXiv:2505.05849. Pass `agent_url` + `instructed_tool` (+ optional `tasks`/`seed_payload`/`n_iterations`).
- **generate_eva_attack** — Generate + auto-execute an EVA evolving environmental-injection search against a GUI/computer-use agent: seeds a pop-up/overlay and evolves it (Trust/Urgency, k_max iters) with a two-stage action-region + LLM intent evaluator. Our implementation of arXiv:2505.14289. Pass `agent_url` + `instructed_tool`.
Expand Down
2 changes: 1 addition & 1 deletion capabilities/ai-red-teaming/capability.yaml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
schema: 1
name: ai-red-teaming
version: "1.17.5"
version: "1.17.6"
description: >
Probe the security and safety of AI applications, agents, and foundation models.
Orchestrates adversarial attack workflows to discover vulnerabilities in LLMs,
Expand Down
25 changes: 10 additions & 15 deletions capabilities/ai-red-teaming/scripts/attack_runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -5072,17 +5072,14 @@ def _generate_atlas_single(config: dict, agent_config: dict) -> str:
def generate_multistep_tool_attack(params: dict) -> dict:
"""Generate a multi-step tool-attack search against an HTTP agent.

Our own implementation of the algorithmic multi-step tool-attack techniques
surfaced by the Kaggle 'AI Agent Security - Multi-Step Tool Attacks'
competition (Kaggle x OpenAI x Google x IEEE, 2026), generalized to any
HTTP agent. We reimplement the algorithms (not the competition code) so we
control the trace format and success predicate:
- Replay-stable candidates (Pilkwang Kim, 'Replay-Dense Exfiltration')
- Shared-prefix / snapshot branching (competition starter; Trajectory Search)
- Budget-aware N_eff sizing (Pilkwang Kim, replay-budget model)
- Two-probe recovery (tetsu2131)
- Split / multi-post exfiltration (Nguyen Cong Tuan, 'JED multipost')
https://www.kaggle.com/competitions/ai-agent-security-multi-step-tool-attacks
A black-box multi-step tool-attack search, generalized to any HTTP agent. We
own the trace format and success predicate:
- Replay-stable candidates
- Shared-prefix / snapshot branching (trajectory search)
- Budget-aware N_eff sizing (replay-budget model)
- Calibrated two-probe recovery
- Split / multi-post exfiltration
For a white-box target you own, swap in a gradient-guided proposer (GCG / HotFlip).

Requires: agent_url. Optional: secret_source, sink_url, sink_email,
time_budget_s, attacker_model, assessment_name.
Expand Down Expand Up @@ -5140,10 +5137,8 @@ def generate_multistep_tool_attack(params: dict) -> dict:
class MultiStepToolAttack:
"""Budgeted search for a replay-stable read-then-exfiltrate causal path.

Attribution (our own implementation): Kaggle AI Agent Security - Multi-Step
Tool Attacks (Kaggle x OpenAI x Google x IEEE, 2026). Replay-stable
candidates + trajectory/shared-prefix search (Pilkwang Kim); two-probe
recovery (tetsu2131); split/multi-post exfiltration (JED / Nguyen Cong Tuan).
Black-box: replay-stable candidates + trajectory / shared-prefix search, a
calibrated two-probe recovery step, and split / multi-post exfiltration.
"""

def __init__(self, target, secret_source, sink_url, sink_email, time_budget_s=120.0):
Expand Down
82 changes: 23 additions & 59 deletions capabilities/ai-red-teaming/tools/attacks.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@
# loaded as flat modules (no parent package), so relative imports do not work.
import importlib.util as _ilu
from pathlib import Path as _Path

_errors_path = _Path(__file__).resolve().parent / "_errors.py"
_spec = _ilu.spec_from_file_location("airt_tools_errors", _errors_path)
_errors_mod = _ilu.module_from_spec(_spec)
Expand Down Expand Up @@ -316,12 +317,10 @@ def generate_multistep_tool_attack(
path (the agent reads an untrusted source, then performs a privileged
outbound action carrying its contents), verified by a trace predicate.

Our own implementation of the algorithmic techniques from the Kaggle
'AI Agent Security - Multi-Step Tool Attacks' competition (Kaggle x OpenAI x
Google x IEEE, 2026): replay-stable candidates + trajectory/shared-prefix
search (Pilkwang Kim), two-probe recovery (tetsu2131), and split/multi-post
exfiltration (JED / Nguyen Cong Tuan). Generalized to any agent/tools.
https://www.kaggle.com/competitions/ai-agent-security-multi-step-tool-attacks
Algorithmic techniques: replay-stable candidates, trajectory / shared-prefix
search, a calibrated two-probe recovery step, and split / multi-post
exfiltration. Black-box and generalized to any agent/tools; for a white-box
target, swap in a gradient-guided proposer (GCG / HotFlip).
"""
params: dict[str, t.Any] = {
"agent_url": agent_url,
Expand All @@ -341,9 +340,7 @@ def generate_multistep_tool_attack(
def generate_agentvigil_attack(
agent_url: t.Annotated[str, "HTTP endpoint of the tool-using agent to red-team"],
instructed_tool: t.Annotated[str, "Tool whose firing counts as injection success"] = "send_email",
tasks: t.Annotated[
list[str] | None, "Benign tasks to compute ASR over; payload is planted in each"
] = None,
tasks: t.Annotated[list[str] | None, "Benign tasks to compute ASR over; payload is planted in each"] = None,
seed_payload: t.Annotated[str, "Initial injection payload (MCTS root)"] = "",
attacker_model: t.Annotated[str, "LLM that mutates payloads (five operators)"] = "dn/claude-opus-4-8",
n_iterations: t.Annotated[int, "MCTS rollouts"] = 30,
Expand Down Expand Up @@ -621,15 +618,9 @@ def generate_extraction_attack(
] = "knockoff",
api_url: t.Annotated[str, "Target classifier predict endpoint (POST)."] = "",
api_key: t.Annotated[str, "API key for the x-api-key header (optional)."] = "",
pool_url: t.Annotated[
str, "GET endpoint returning {inputs: [...]} — the unlabeled query pool."
] = "",
query_pool: t.Annotated[
list | None, "Inline query pool (used if pool_url is not given)."
] = None,
request_template: t.Annotated[
str, "Request body with a single {input} placeholder."
] = '{"features": {input}}',
pool_url: t.Annotated[str, "GET endpoint returning {inputs: [...]} — the unlabeled query pool."] = "",
query_pool: t.Annotated[list | None, "Inline query pool (used if pool_url is not given)."] = None,
request_template: t.Annotated[str, "Request body with a single {input} placeholder."] = '{"features": {input}}',
probabilities_path: t.Annotated[str, "JSONPath to the probability vector."] = "$.probabilities",
input_format: t.Annotated[str, "json_array | image_b64 | text."] = "json_array",
num_classes: t.Annotated[int, "Number of classes."] = 2,
Expand Down Expand Up @@ -682,15 +673,11 @@ def generate_membership_attack(
] = "threshold",
api_url: t.Annotated[str, "Target classifier predict endpoint (POST)."] = "",
api_key: t.Annotated[str, "API key for the x-api-key header (optional)."] = "",
members_url: t.Annotated[
str, "GET endpoint returning {records: [...], labels: [...]} for training members."
] = "",
members_url: t.Annotated[str, "GET endpoint returning {records: [...], labels: [...]} for training members."] = "",
nonmembers_url: t.Annotated[str, "GET endpoint for held-out non-members."] = "",
members: t.Annotated[list | None, "Inline member records (if no members_url)."] = None,
nonmembers: t.Annotated[list | None, "Inline non-member records (if no nonmembers_url)."] = None,
request_template: t.Annotated[
str, "Request body with a single {input} placeholder."
] = '{"features": {input}}',
request_template: t.Annotated[str, "Request body with a single {input} placeholder."] = '{"features": {input}}',
probabilities_path: t.Annotated[str, "JSONPath to the probability vector."] = "$.probabilities",
input_format: t.Annotated[str, "json_array | image_b64 | text."] = "json_array",
num_classes: t.Annotated[int, "Number of classes."] = 2,
Expand Down Expand Up @@ -731,25 +718,17 @@ def generate_membership_attack(

@safe_tool
def generate_inversion_attack(
attack_type: t.Annotated[
str, "Model-inversion attack: confidence (MI-Face hill-climb) or nes."
] = "confidence",
attack_type: t.Annotated[str, "Model-inversion attack: confidence (MI-Face hill-climb) or nes."] = "confidence",
api_url: t.Annotated[str, "Target classifier predict endpoint (POST)."] = "",
api_key: t.Annotated[str, "API key for the x-api-key header (optional)."] = "",
num_classes: t.Annotated[int, "Number of classes."] = 2,
input_dim: t.Annotated[
int, "Feature-vector length (tabular). Inferred from the target's /pool if omitted."
] = 0,
input_dim: t.Annotated[int, "Feature-vector length (tabular). Inferred from the target's /pool if omitted."] = 0,
input_shape: t.Annotated[
str, "Image shape as 'H,W' (e.g. '8,8'). Inferred from /pool when square, if omitted."
] = "",
target_classes: t.Annotated[
list | None, "Classes to reconstruct (default: all classes)."
] = None,
target_classes: t.Annotated[list | None, "Classes to reconstruct (default: all classes)."] = None,
max_queries: t.Annotated[int, "Max target queries."] = 1500,
request_template: t.Annotated[
str, "Request body with a single {input} placeholder."
] = '{"features": {input}}',
request_template: t.Annotated[str, "Request body with a single {input} placeholder."] = '{"features": {input}}',
probabilities_path: t.Annotated[str, "JSONPath to the probability vector."] = "$.probabilities",
input_format: t.Annotated[str, "json_array | image_b64 | text."] = "json_array",
modality: t.Annotated[str, "tabular | image | text."] = "tabular",
Expand Down Expand Up @@ -794,15 +773,9 @@ def generate_evasion_attack(
] = "boundary",
api_url: t.Annotated[str, "Target classifier predict endpoint (POST)."] = "",
api_key: t.Annotated[str, "API key for the x-api-key header (optional)."] = "",
sample_url: t.Annotated[
str, "GET endpoint returning {inputs: [...]} - the first input is perturbed."
] = "",
original: t.Annotated[
object, "Inline original input to perturb (if no sample_url)."
] = None,
request_template: t.Annotated[
str, "Request body with a single {input} placeholder."
] = '{"features": {input}}',
sample_url: t.Annotated[str, "GET endpoint returning {inputs: [...]} - the first input is perturbed."] = "",
original: t.Annotated[object, "Inline original input to perturb (if no sample_url)."] = None,
request_template: t.Annotated[str, "Request body with a single {input} placeholder."] = '{"features": {input}}',
probabilities_path: t.Annotated[str, "JSONPath to the probability vector."] = "$.probabilities",
input_format: t.Annotated[str, "json_array | image_b64 | text."] = "json_array",
num_classes: t.Annotated[int, "Number of classes."] = 2,
Expand Down Expand Up @@ -933,9 +906,7 @@ def generate_multimodal_attack(
custom_region: t.Annotated[
str, "AWS region for a streaming (Nova Sonic) or aws_sigv4 target (default us-east-1)."
] = "",
custom_service: t.Annotated[
str, "AWS service for an aws_sigv4 HTTP target (default 'sagemaker')."
] = "",
custom_service: t.Annotated[str, "AWS service for an aws_sigv4 HTTP target (default 'sagemaker')."] = "",
custom_request_format: t.Annotated[
str,
"Request body encoding for custom_url: 'json' (default, renders the request "
Expand All @@ -947,9 +918,7 @@ def generate_multimodal_attack(
] = "",
custom_voice: t.Annotated[str, "Voice id for a Nova Sonic streaming target (default matthew)."] = "",
custom_system_prompt: t.Annotated[str, "System prompt for a streaming S2S target."] = "",
custom_model_id: t.Annotated[
str, "Model id for a streaming target (default amazon.nova-sonic-v1:0)."
] = "",
custom_model_id: t.Annotated[str, "Model id for a streaming target (default amazon.nova-sonic-v1:0)."] = "",
score_media_output: t.Annotated[
bool,
"Score the target's GENERATED media (image-out / speech-to-speech), not just its "
Expand Down Expand Up @@ -1139,8 +1108,7 @@ def generate_multimodal_category_attack(
def build_media_manifest(
directory: t.Annotated[
str,
"Directory of media files to inventory (recursively). Use this for "
'"the images in ./imgs" style requests.',
"Directory of media files to inventory (recursively). Use this for " '"the images in ./imgs" style requests.',
] = "",
paths: t.Annotated[
list[str] | None,
Expand Down Expand Up @@ -1196,12 +1164,8 @@ def generate_injection_images(
"Path to a CSV whose first column is the attack text (one image per row). "
"Use when the user hands you a CSV of prompts to render as images.",
] = "",
output_dir: t.Annotated[
str, "Directory to write the images into (default ./injection_images)."
] = "",
base_image: t.Annotated[
str, "Optional base image path to overlay the text onto (else a plain canvas)."
] = "",
output_dir: t.Annotated[str, "Directory to write the images into (default ./injection_images)."] = "",
base_image: t.Annotated[str, "Optional base image path to overlay the text onto (else a plain canvas)."] = "",
font_size: t.Annotated[int, "Font size for the rendered text (default 40)."] = 40,
) -> str:
"""Render attack text into typographic/visual prompt-injection IMAGES.
Expand Down
Loading