Skip to content
FSoft-AI4CodePublic

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

20 Commits

Folders and files

Repository files navigation

Attacker-Verifier

Security feedback for coding agents, grounded in real attacks and their execution traces.

Claude Code plugin version 0.2.0 python 3.8+ license MIT

Quick start · How it works · Benchmark results · Experiments · Configuration · Paper

Overview: the coding agent's patch is attacked with PoC probes, unfaithful probes are rejected, the admitted probes run in a sandbox, and the verifier turns their traces into a verdict and feedback.


Attacker-Verifier checks code by attacking it. An attacker writes small proof-of-concept probes that call the code with adversarial inputs and never say what the output should be. Probes that fake their evidence are thrown out; the rest run in a sandbox, and a verifier decides from each execution trace whether the code did something unsafe. A reported violation always comes with the probe and the trace that show it, so a coding agent can fix the cause instead of guessing.

The method is FALCON, from the paper Secure Agentic Coding through Counterexample-Grounded Feedback. This repository has FALCON as a Claude Code plugin (named attacker-verifier) and the code for the paper's experiments. The figures and tables below come from the paper.

Quick start

In Claude Code:

/plugin marketplace add SafeCodeAgent/FALCON
/plugin install attacker-verifier@attacker-verifier

Audit existing code, or write new code and harden it:

/check-code-security src/storage.py
/secure-code-generation add an endpoint that serves report files by name

/check-code-security writes a report named <repo>_<timestamp>.md (example). /secure-code-generation implements the request, then attacks and repairs it for two rounds by default. The first run in a session asks whether probes should run in Docker or on the host. The probes need Python 3.8 or newer; the engine uses only the standard library.

How it works

Prior approaches ask a model to reason about code or to define expected outcomes; Attacker-Verifier separates exploration from verification through execution evidence.

A judge that reads source code has to infer how the program behaves, and a model that writes security tests also has to decide the correct output. Mistakes in either let a candidate pass while keeping the vulnerability. Attacker-Verifier splits the job in two and connects the halves through execution:

  1. Targets. The changed functions, methods, and classes in non-test files are selected (the plugin can also rank a whole repository by security-relevant code).
  2. Attack. An attacker writes deterministic probes that call each target with adversarial inputs or a prepared environment. A probe records what happened and never states the expected result.
  3. Faithfulness. Before running, a probe is rejected if it redefines, rebinds, or patches the target, writes code into the repository, runs the sensitive operation itself, or prints evidence it made up. After running, a probe that never reached the target, or whose output cannot be tied to it, is discarded.
  4. Execution. Each probe runs alone in a sandbox, under a tracer that records the target's calls, arguments, return values, exceptions, and side effects.
  5. Verification. Deterministic checks decide traces with clear runtime evidence, such as a crash on the target path. A judge model reads the remaining traces, never the source, and decides whether the observed behaviour is a violation.
  6. Feedback. A target is insecure if any admitted probe is insecure, secure if at least one probe was admitted and none is insecure, and has no evidence otherwise. Each violation goes back to the coding agent with its probe and trace.

Benchmark results

All numbers are from the paper. Func-Sec@1 counts a task only when both the benchmark's functional tests and its security tests pass; those tests are never shown to the attacker, the verifier, or the coding model.

Reported progress versus held-out security

Repair on CWEval with four security signals: held-out Func-Sec@1 per round, the improvement each signal reports, and reported versus held-out improvement.

GPT-5.4-mini repaired its CWEval solutions for five rounds against four different signals. Every signal reported progress on its own verdicts, but only part of it reached the held-out tests: Attacker-Verifier raised Func-Sec@1 by 25.2 points (61.3 to 86.6), CodeQL by 6.7, the LLM judge by 0.0, and attacker-written tests lowered it by 1.7.

Test-time repair

Func-Sec@1 on SecCodeBench-V2 before and after repair, with frontier models as single-pass references. Func-Sec@1 on CWEval before and after repair, with frontier models as single-pass references.

Up to five rounds of repair on SecCodeBench-V2 and CWEval, with the attacker and the judge on the same model as the coder (starred). Func-Sec@1 rises by 22.5 to 24.0 points on SecCodeBench-V2 and 24.4 to 25.2 points on CWEval. GPT-5.4-nano goes from 40.6% to 64.3% and from 52.9% to 77.3%, close to or above Claude Opus 4.8 without repair.

Functionality, security, and Func-Sec per repair round

Func, Sec, and Func-Sec over five repair rounds on SecCodeBench-V2 and CWEval.

Coding agents on SWE-bench-style secure code generation tasks

Three coding-agent harnesses, SWE-agent, Claude Code, and OpenCode, edit real Python repositories in a SWE-bench-style secure-coding benchmark (SusVibes, 186 tasks across 77 CWEs). Before a patch is accepted, the changed code is attacked and any violation goes back to the agent, for at most five checks; the attacker and the judge use the agent's model (DeepSeek-V4-Pro is the 08-13 release). FuncPass and SecPass are the benchmark's task-level criteria; Func-TestRate and Sec-TestRate are the fractions of functional and security tests passed. Values are averaged over the three harnesses.

Model FuncPass SecPass Func‑TestRate Sec‑TestRate
GPT‑5.4‑mini 44.09 12.36 74.01 69.76
  + FALCON

46.59$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{2.50}}$

17.20$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{4.84}}$

76.24$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{2.23}}$

72.55$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{2.79}}$

GPT‑5.4 68.82 20.25 81.65 77.55
  + FALCON

70.97$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{2.15}}$

25.27$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{5.02}}$

83.12$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{1.47}}$

79.35$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{1.80}}$

GPT‑5.6‑luna 62.54 19.00 80.99 77.78
  + FALCON

63.62$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{1.08}}$

22.22$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{3.22}}$

81.47$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{0.48}}$

79.36$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{1.58}}$

DeepSeek‑V4‑Pro 81.90 20.79 83.87 79.92
  + FALCON

82.61$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{0.71}}$

24.91$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{4.12}}$

84.68$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{0.81}}$

81.59$\enspace\color{#1f9e46}{\scriptstyle\blacktriangle\thinspace\mathsf{1.67}}$

Reinforcement learning

GRPO training reward, held-out Func-Sec@1, and KL for six security rewards on Qwen2.5-Coder-3B and 7B.

GRPO on SecCodePLT+ with Qwen2.5-Coder 3B and 7B, six security rewards, and the same functionality reward. Every policy raises its own training reward, but the static-analysis (REAL) and learned (SecCodePRM) rewards transfer much less to held-out Func-Sec@1, and at 3B it falls while their reward keeps rising. Adding the Attacker-Verifier reward to either one gives the best held-out results (57.25 at 3B with REAL, 77.26 at 7B with SecCodePRM); combining REAL with SecCodePRM does not help.

Verifier accuracy and cost

Agreement with CWEval's labeled security tests, averaged over code from three models. Judging one observed execution is easier than judging all the executions a piece of code allows, and it is cheaper.

Verifier (judge: GPT-5.2) Accuracy False positives False negatives Cost per example
Deterministic checks only 97.54% 2.33% 2.62% $0
Deterministic checks + trace judge 98.58% 2.73% 0.20% $0.085
Trace judge on every trace 87.99% 23.92% 0.18% $0.634
Judge reading the source code 61.16% 50.43% 28.05% $1.227
CodeQL 56.68% 17.95% 66.42% $0

Using the plugin

Commands

/check-code-security                       # asks: whole repository or a part you name
/check-code-security src/ --probes 8-12
/check-code-security --scope changed       # uncommitted work, including new files
/secure-code-generation implement the CSV import --turns 3

The report lists each unsafe function with the weakness, the reason, the evidence, and a trace excerpt, followed by per-target results and the probes the faithfulness check removed.

Sandbox

The first time either command runs in a session, it asks how to run probes: in a Docker image (a fresh container per probe) or a running container, with the repository mounted inside it, or on the host with the permissions the Claude Code session already has. The answer is kept for the session. See docs/sandbox.md.

Configuration

Every setting has a default. Change them in /config (or /plugin, then attacker-verifier, then Configure), or for one run with arguments:

Setting Default Argument
Attacker model main coding agent --attacker NAME
Verifier model (trace judge) main coding agent --verifier NAME
Attacker / verifier reasoning effort inherit --attacker-effort, --verifier-effort
Attacker max turns per target 20 --attacker-max-turns N
Probes per target 5 to 10 --probes MIN-MAX
Repair rounds (/secure-code-generation) 2 --turns N
Max targets per run 20 --max-targets N

main coding agent runs the role in your current session; a model name runs it as a subagent on that model. A repository can pin engine settings in .attacker-verifier/config.json. docs/configuration.md lists everything.

Limits

  • Only Python is attacked. Files in other languages are listed in the report as not attacked.
  • A secure result means the probes that ran found nothing, not that the code is proven safe. A larger probe budget or a wider scope checks more.
  • Running code takes longer than reading it.

Reproducing the experiments

Experiment Paper Code
Test-time repair on CWEval and SecCodeBench-V2 §4.2.1–4.2.2 experiments/test-time-repair
Coding agents on SusVibes §4.2.3 experiments/susvibes
GRPO on SecCodePLT+ §4.2.4 experiments/rl

Benchmarks, task images, and model weights are not included, and the held-out tests run only in the separate grading commands. See experiments/README.md.

Repository layout

.claude-plugin/   plugin and marketplace manifests
skills/           /check-code-security and /secure-code-generation
agents/           attacker and verifier subagents
engine/           target selection, faithfulness checks, sandbox, crash oracle, report
docs/             procedures, configuration reference, sample report
tests/            engine tests (python3 -m pytest tests)
experiments/      code for the paper's experiments
img/              logo and figures

License

MIT, see LICENSE. The experiment code includes third-party components under their own licenses; see experiments/README.md.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages