Quick start · How it works · Benchmark results · Experiments · Configuration · Paper
Attacker-Verifier checks code by attacking it. An attacker writes small proof-of-concept probes that call the code with adversarial inputs and never say what the output should be. Probes that fake their evidence are thrown out; the rest run in a sandbox, and a verifier decides from each execution trace whether the code did something unsafe. A reported violation always comes with the probe and the trace that show it, so a coding agent can fix the cause instead of guessing.
The method is FALCON, from the paper Secure Agentic Coding through Counterexample-Grounded Feedback. This repository has FALCON as a Claude Code plugin (named attacker-verifier) and the code for the paper's experiments. The figures and tables below come from the paper.
In Claude Code:
/plugin marketplace add SafeCodeAgent/FALCON
/plugin install attacker-verifier@attacker-verifier
Audit existing code, or write new code and harden it:
/check-code-security src/storage.py
/secure-code-generation add an endpoint that serves report files by name
/check-code-security writes a report named <repo>_<timestamp>.md
(example). /secure-code-generation implements the
request, then attacks and repairs it for two rounds by default. The first run in
a session asks whether probes should run in Docker or on the host. The probes
need Python 3.8 or newer; the engine uses only the standard library.
A judge that reads source code has to infer how the program behaves, and a model that writes security tests also has to decide the correct output. Mistakes in either let a candidate pass while keeping the vulnerability. Attacker-Verifier splits the job in two and connects the halves through execution:
- Targets. The changed functions, methods, and classes in non-test files are selected (the plugin can also rank a whole repository by security-relevant code).
- Attack. An attacker writes deterministic probes that call each target with adversarial inputs or a prepared environment. A probe records what happened and never states the expected result.
- Faithfulness. Before running, a probe is rejected if it redefines, rebinds, or patches the target, writes code into the repository, runs the sensitive operation itself, or prints evidence it made up. After running, a probe that never reached the target, or whose output cannot be tied to it, is discarded.
- Execution. Each probe runs alone in a sandbox, under a tracer that records the target's calls, arguments, return values, exceptions, and side effects.
- Verification. Deterministic checks decide traces with clear runtime evidence, such as a crash on the target path. A judge model reads the remaining traces, never the source, and decides whether the observed behaviour is a violation.
- Feedback. A target is insecure if any admitted probe is insecure, secure if at least one probe was admitted and none is insecure, and has no evidence otherwise. Each violation goes back to the coding agent with its probe and trace.
All numbers are from the paper. Func-Sec@1 counts a task only when both the benchmark's functional tests and its security tests pass; those tests are never shown to the attacker, the verifier, or the coding model.
GPT-5.4-mini repaired its CWEval solutions for five rounds against four different signals. Every signal reported progress on its own verdicts, but only part of it reached the held-out tests: Attacker-Verifier raised Func-Sec@1 by 25.2 points (61.3 to 86.6), CodeQL by 6.7, the LLM judge by 0.0, and attacker-written tests lowered it by 1.7.
Up to five rounds of repair on SecCodeBench-V2 and CWEval, with the attacker and the judge on the same model as the coder (starred). Func-Sec@1 rises by 22.5 to 24.0 points on SecCodeBench-V2 and 24.4 to 25.2 points on CWEval. GPT-5.4-nano goes from 40.6% to 64.3% and from 52.9% to 77.3%, close to or above Claude Opus 4.8 without repair.
Three coding-agent harnesses, SWE-agent, Claude Code, and OpenCode, edit real Python repositories in a SWE-bench-style secure-coding benchmark (SusVibes, 186 tasks across 77 CWEs). Before a patch is accepted, the changed code is attacked and any violation goes back to the agent, for at most five checks; the attacker and the judge use the agent's model (DeepSeek-V4-Pro is the 08-13 release). FuncPass and SecPass are the benchmark's task-level criteria; Func-TestRate and Sec-TestRate are the fractions of functional and security tests passed. Values are averaged over the three harnesses.
| Model | FuncPass | SecPass | Func‑TestRate | Sec‑TestRate |
|---|---|---|---|---|
| GPT‑5.4‑mini | 44.09 | 12.36 | 74.01 | 69.76 |
| + FALCON |
46.59 |
17.20 |
76.24 |
72.55 |
| GPT‑5.4 | 68.82 | 20.25 | 81.65 | 77.55 |
| + FALCON |
70.97 |
25.27 |
83.12 |
79.35 |
| GPT‑5.6‑luna | 62.54 | 19.00 | 80.99 | 77.78 |
| + FALCON |
63.62 |
22.22 |
81.47 |
79.36 |
| DeepSeek‑V4‑Pro | 81.90 | 20.79 | 83.87 | 79.92 |
| + FALCON |
82.61 |
24.91 |
84.68 |
81.59 |
GRPO on SecCodePLT+ with Qwen2.5-Coder 3B and 7B, six security rewards, and the same functionality reward. Every policy raises its own training reward, but the static-analysis (REAL) and learned (SecCodePRM) rewards transfer much less to held-out Func-Sec@1, and at 3B it falls while their reward keeps rising. Adding the Attacker-Verifier reward to either one gives the best held-out results (57.25 at 3B with REAL, 77.26 at 7B with SecCodePRM); combining REAL with SecCodePRM does not help.
Agreement with CWEval's labeled security tests, averaged over code from three models. Judging one observed execution is easier than judging all the executions a piece of code allows, and it is cheaper.
| Verifier (judge: GPT-5.2) | Accuracy | False positives | False negatives | Cost per example |
|---|---|---|---|---|
| Deterministic checks only | 97.54% | 2.33% | 2.62% | $0 |
| Deterministic checks + trace judge | 98.58% | 2.73% | 0.20% | $0.085 |
| Trace judge on every trace | 87.99% | 23.92% | 0.18% | $0.634 |
| Judge reading the source code | 61.16% | 50.43% | 28.05% | $1.227 |
| CodeQL | 56.68% | 17.95% | 66.42% | $0 |
/check-code-security # asks: whole repository or a part you name
/check-code-security src/ --probes 8-12
/check-code-security --scope changed # uncommitted work, including new files
/secure-code-generation implement the CSV import --turns 3
The report lists each unsafe function with the weakness, the reason, the evidence, and a trace excerpt, followed by per-target results and the probes the faithfulness check removed.
The first time either command runs in a session, it asks how to run probes: in a Docker image (a fresh container per probe) or a running container, with the repository mounted inside it, or on the host with the permissions the Claude Code session already has. The answer is kept for the session. See docs/sandbox.md.
Every setting has a default. Change them in /config (or /plugin, then
attacker-verifier, then Configure), or for one run with arguments:
| Setting | Default | Argument |
|---|---|---|
| Attacker model | main coding agent |
--attacker NAME |
| Verifier model (trace judge) | main coding agent |
--verifier NAME |
| Attacker / verifier reasoning effort | inherit |
--attacker-effort, --verifier-effort |
| Attacker max turns per target | 20 | --attacker-max-turns N |
| Probes per target | 5 to 10 | --probes MIN-MAX |
Repair rounds (/secure-code-generation) |
2 | --turns N |
| Max targets per run | 20 | --max-targets N |
main coding agent runs the role in your current session; a model name runs it
as a subagent on that model. A repository can pin engine settings in
.attacker-verifier/config.json. docs/configuration.md
lists everything.
- Only Python is attacked. Files in other languages are listed in the report as not attacked.
- A secure result means the probes that ran found nothing, not that the code is proven safe. A larger probe budget or a wider scope checks more.
- Running code takes longer than reading it.
| Experiment | Paper | Code |
|---|---|---|
| Test-time repair on CWEval and SecCodeBench-V2 | §4.2.1–4.2.2 | experiments/test-time-repair |
| Coding agents on SusVibes | §4.2.3 | experiments/susvibes |
| GRPO on SecCodePLT+ | §4.2.4 | experiments/rl |
Benchmarks, task images, and model weights are not included, and the held-out tests run only in the separate grading commands. See experiments/README.md.
.claude-plugin/ plugin and marketplace manifests
skills/ /check-code-security and /secure-code-generation
agents/ attacker and verifier subagents
engine/ target selection, faithfulness checks, sandbox, crash oracle, report
docs/ procedures, configuration reference, sample report
tests/ engine tests (python3 -m pytest tests)
experiments/ code for the paper's experiments
img/ logo and figures
MIT, see LICENSE. The experiment code includes third-party components under their own licenses; see experiments/README.md.






