Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,14 @@
results and metadata, and rejects non-JSON or empty output; the console shows the decoded words
while the agent runs. `{{packageDir}}` in a command names the installed package directory, for
files that ship with it.
- Claude, Droid, Antigravity and OpenCode now resume their own review conversation
when replying to verdicts, instead of starting fresh with their findings quoted
back. Eight of eleven agents now resume. Each names an explicit session id —
Claude via an assigned `--session-id`, the others read from structured output —
so a reply can never land in an unrelated conversation.
- Droid and Antigravity are read through their JSON output modes, which is where
each reports its session id. Antigravity's reported status is now checked, so a
run that ends early cannot be read as a sign-off.

## 0.6.0 — 2026-09-16

Expand Down
30 changes: 29 additions & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,6 +115,34 @@ your environment. Run records include prompts and outputs, so redact them before

### Authentication, versions, and conversations

A reviewer's reply is a turn in the conversation that raised the findings, so
the agent can see what it said. Every resume names an explicit session id —
never "the latest session", which would answer whatever ran most recently in
that worktree rather than this review.

An id reaches Jury one of two ways. Some CLIs accept one we generate
(`--session-id`), so the conversation is identified before it exists. The rest
print one in structured output, which is read back with `resume.idFrom`. An id
is never inferred from assistant prose: a model that happens to mention a UUID
is not reporting its session, and resuming on that would deliver a verdict into
an unrelated conversation.

| Agent | Session id | Resumes |
|---|---|---|
| Claude | assigned `--session-id` | yes |
| Grok | assigned `--session-id` | yes |
| Qwen | assigned `--session-id` | yes |
| Codex | printed by `exec` | yes |
| Droid | `session_id` in `-o json` | yes |
| Antigravity | `conversation_id` in `--output-format json` | yes |
| OpenCode | `sessionID` in the JSON event stream | yes |
| Kimi Code | printed at the end of the `stream-json` output | yes |
| Amp, Cursor, Copilot | — | no; replies start fresh with the finding context |

Agents that do not resume lose nothing in substance: `buildReply` quotes their
own prior findings back to them. The judge never resumes at all — each finding
is triaged in its own session so one verdict cannot anchor the next.

- Kimi Code: verified CLI 0.39.1 and 0.42.0 (`npm install -g @moonshot-ai/kimi-code`); run `kimi login`, or put an API key in `~/.kimi-code/config.toml`, whose `default_model` is the model used. Prompt mode (`kimi -p`) rejects `--plan`, `--yolo` and `--auto` and approves every tool call itself, so the reviewer boundary is the shipped agent profile: it wraps Kimi's default prompt, removes the Edit and Write tools, and allows only the read-only `explore` sub-agent (the default `coder` sub-agent can write). The shell stays available, so the profile is not a sandbox. The report is read from `--output-format stream-json`, and replies resume the exact session id Kimi prints at the end of the stream; the profile flag is omitted on resume because Kimi refuses it next to `--session`, and the session keeps the agent it was created with. Kimi keeps its own session history per working directory under `~/.kimi-code`. The legacy Python `kimi-cli` installs an executable of the same name but takes different flags (`kimi --version` prints 1.x for it and 0.x for Kimi Code); to keep using it, override the entry in `jury.config.json`:

```json
Expand All @@ -123,6 +151,6 @@ your environment. Run records include prompts and outputs, so redact them before
- Cursor: verified `cursor-agent` 2025.10.28-0a91dc2; run `cursor-agent login`. The generic executable `agent` may belong to another product, so Jury uses `cursor-agent`. Final JSON must explicitly report success. Replies start fresh.
- Copilot: verified CLI 0.0.392 flags; authenticate with interactive `/login`. A reviewer can inspect files and Git but cannot use write tools; the judge allows tools. Local permission configuration remains trusted. Replies start fresh.
- Qwen: targets CLI 0.23.4 (`npm install -g @qwen-code/qwen-code`); run `qwen` and complete `/auth` before headless use. It runs in the process worktree, not merely an added access directory. Reviews use assigned UUIDs, and replies resume only that UUID; no implicit latest session. Final JSON must explicitly report success.
- Amp: verified CLI 0.0.1788739286; run `amp login`. Threads are private and IDE context is disabled. Both roles can execute tools automatically; replies start fresh with their own finding context.
- Amp: verified CLI 0.0.1788739286; run `amp login`. Threads are private and IDE context is disabled. Both roles can execute tools automatically; replies start fresh with their own finding context. (`--stream-json` does expose a thread id, so resume is possible once that output mode is parsed.)

A missing executable, nonzero exit, timeout, or unsuccessful structured result never approves a review. Git command allowlists are CLI tool permissions, not OS sandboxes; use trusted local agent configuration. Model credentials and service availability are prerequisites for live inference.
2 changes: 1 addition & 1 deletion lib/agent-schema.js
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ export const PROMPT_DELIVERY = ["argv", "file", "stdin"];
export const CWD_MODES = ["worktree", "flag"];

/** How much of stdout is the agent's report. */
export const REPORT_MODES = ["whole", "tail", "opencode-json", "result-json", "kimi-json"];
export const REPORT_MODES = ["whole", "tail", "opencode-json", "result-json", "kimi-json", "agy-json"];

/**
* How strongly the reviewer invocation is prevented from writing.
Expand Down
14 changes: 14 additions & 0 deletions lib/agents.js
Original file line number Diff line number Diff line change
Expand Up @@ -175,6 +175,20 @@ export function extractReport(stdout, mode = "whole") {
}
return result.result.trim();
}
if (mode === "agy-json") {
let data;
try { data = JSON.parse(text); } catch { throw new Error("Invalid Antigravity JSON output"); }
// Antigravity reports its own outcome. A timeout that still exits 0 with a
// partial answer arrives as a non-SUCCESS status, and reading the prose
// regardless is how a truncated review gets accepted as a sign-off.
if (data?.status !== "SUCCESS") {
throw new Error(`Antigravity did not finish: ${data?.status ?? "missing status"}`);
}
if (typeof data.response !== "string" || !data.response.trim()) {
throw new Error("Antigravity returned no response text");
}
return data.response.trim();
}
if (mode === "opencode-json") {
let parts = [];
let messageId;
Expand Down
33 changes: 29 additions & 4 deletions lib/agents/agy.json
Original file line number Diff line number Diff line change
Expand Up @@ -5,13 +5,38 @@
"promptDelivery": "argv",
"cwd": "worktree",
"argv": [
"agy", "--dangerously-skip-permissions", "--add-dir", "{{worktree}}",
"--print-timeout", "20m", "--print", "{{promptText}}"
"agy",
"--dangerously-skip-permissions",
"--add-dir",
"{{worktree}}",
"--print-timeout",
"20m",
"--output-format",
"json",
"--print",
"{{promptText}}"
],
"sandbox": "none",
"sandboxNote": "--dangerously-skip-permissions approves every tool, including writes",
"resume": { "supported": false, "reason": "print mode starts a fresh session per run" },
"report": "whole",
"resume": {
"supported": true,
"argv": [
"agy",
"--dangerously-skip-permissions",
"--add-dir",
"{{worktree}}",
"--print-timeout",
"20m",
"--output-format",
"json",
"--conversation",
"{{sessionId}}",
"--print",
"{{promptText}}"
],
"idFrom": "\"conversation_id\"\\s*:\\s*\"([^\"]+)\""
},
"report": "agy-json",
"expectSeconds": 420,
"install": "npm install -g @google/antigravity-cli",
"docs": "https://antigravity.google/docs/cli"
Expand Down
45 changes: 36 additions & 9 deletions lib/agents/claude.json
Original file line number Diff line number Diff line change
Expand Up @@ -5,21 +5,48 @@
"promptDelivery": "argv",
"cwd": "worktree",
"argv": [
"claude", "-p", "{{promptText}}",
"--permission-mode", "plan",
"--disallowedTools", "Edit,Write,NotebookEdit",
"--add-dir", "{{worktree}}"
"claude",
"-p",
"{{promptText}}",
"--session-id",
"{{sessionId}}",
"--permission-mode",
"plan",
"--disallowedTools",
"Edit,Write,NotebookEdit",
"--add-dir",
"{{worktree}}"
],
"judgeArgv": [
"claude", "-p", "{{promptText}}",
"--permission-mode", "acceptEdits",
"--add-dir", "{{worktree}}"
"claude",
"-p",
"{{promptText}}",
"--permission-mode",
"acceptEdits",
"--add-dir",
"{{worktree}}"
],
"sandbox": "plan",
"sandboxNote": "reviews run in plan mode with Edit, Write and NotebookEdit withheld",
"resume": { "supported": false, "reason": "each finding is judged on its own merits" },
"resume": {
"supported": true,
"argv": [
"claude",
"-p",
"{{promptText}}",
"--resume",
"{{sessionId}}",
"--permission-mode",
"plan",
"--disallowedTools",
"Edit,Write,NotebookEdit",
"--add-dir",
"{{worktree}}"
]
},
"report": "whole",
"expectSeconds": 900,
"install": "npm install -g @anthropic-ai/claude-code",
"docs": "https://docs.claude.com/en/docs/claude-code/cli-reference"
"docs": "https://docs.claude.com/en/docs/claude-code/cli-reference",
"newSession": true
}
34 changes: 31 additions & 3 deletions lib/agents/droid.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,11 +4,39 @@
"role": "reviewer",
"promptDelivery": "file",
"cwd": "flag",
"argv": ["droid", "exec", "--cwd", "{{worktree}}", "--auto", "medium", "-f", "{{promptFile}}"],
"argv": [
"droid",
"exec",
"--cwd",
"{{worktree}}",
"--auto",
"medium",
"-o",
"json",
"-f",
"{{promptFile}}"
],
"sandbox": "plan",
"sandboxNote": "--auto medium withholds destructive actions but permits edits",
"resume": { "supported": false, "reason": "exec output carries no session id" },
"report": "whole",
"resume": {
"supported": true,
"argv": [
"droid",
"exec",
"--cwd",
"{{worktree}}",
"--auto",
"medium",
"-o",
"json",
"-s",
"{{sessionId}}",
"-f",
"{{promptFile}}"
],
"idFrom": "\"session_id\"\\s*:\\s*\"([^\"]+)\""
},
"report": "result-json",
"expectSeconds": 130,
"install": "curl -fsSL https://app.factory.ai/cli | sh",
"docs": "https://docs.factory.ai/cli/getting-started/quickstart"
Expand Down
18 changes: 16 additions & 2 deletions lib/agents/opencode.json
Original file line number Diff line number Diff line change
Expand Up @@ -38,8 +38,22 @@
"sandbox": "tools",
"sandboxNote": "OPENCODE_PERMISSION denies edit, task and all bash but a read-only git allowlist",
"resume": {
"supported": false,
"reason": "fresh conversation with finding context; never resume the latest session"
"supported": true,
"argv": [
"opencode",
"run",
"--dir",
"{{worktree}}",
"--agent",
"plan",
"--format",
"json",
"--session",
"{{sessionId}}",
"--",
"{{promptText}}"
],
"idFrom": "\"sessionID\"\\s*:\\s*\"([^\"]+)\""
},
"report": "opencode-json",
"expectSeconds": 600,
Expand Down
9 changes: 7 additions & 2 deletions test/agy.test.js
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,9 @@ test('Antigravity is discoverable without configuration and receives literal pro
const bin = path.join(dir, 'bin');
await mkdir(worktree); await mkdir(bin);
const executable = path.join(bin, 'agy');
await writeFile(executable, `#!${process.execPath}\nimport('node:fs').then(fs => { fs.writeFileSync('invocation.json', JSON.stringify({cwd:process.cwd(), args:process.argv.slice(2)})); console.log('NO NEW FINDINGS'); });\n`, { mode: 0o755 });
// Antigravity is read through --output-format json, so the stub answers in
// that shape: a bare line of prose would no longer be a valid report.
await writeFile(executable, `#!${process.execPath}\nimport('node:fs').then(fs => { fs.writeFileSync('invocation.json', JSON.stringify({cwd:process.cwd(), args:process.argv.slice(2)})); console.log(JSON.stringify({conversation_id:'af730fe4-e36e-4146-a5c5-ba4fea33a325', status:'SUCCESS', response:'NO NEW FINDINGS'})); });\n`, { mode: 0o755 });
const cfg = await loadConfig(worktree, { globalFile: path.join(dir, 'missing.json') });
const agent = reviewers(cfg).find(a => a.name === 'agy');
assert.ok(agent);
Expand All @@ -29,7 +31,10 @@ test('Antigravity is discoverable without configuration and receives literal pro
assert.equal(result.verdict, 'clean');
const call = JSON.parse(await readFile(path.join(worktree, 'invocation.json')));
assert.equal(call.cwd, await realpath(worktree));
assert.deepEqual(call.args, ['--dangerously-skip-permissions', '--add-dir', worktree, '--print-timeout', '20m', '--print', prompt]);
assert.deepEqual(call.args, ['--dangerously-skip-permissions', '--add-dir', worktree, '--print-timeout', '20m', '--output-format', 'json', '--print', prompt]);
// The conversation id is captured so the reply resumes this exact review
// rather than whatever conversation happens to be most recent.
assert.equal(result.sessionId, 'af730fe4-e36e-4146-a5c5-ba4fea33a325');
const env = { ...process.env, HOME: dir, USERPROFILE: dir, PATH: `${bin}${path.delimiter}${process.env.PATH}` };
const listing = spawnSync(process.execPath, [cli, 'agents'], { cwd: worktree, env, encoding: 'utf8' }).stdout;
assert.match(listing, /ok\s+agy\s+reviewer/);
Expand Down
37 changes: 33 additions & 4 deletions test/builtin-agents.test.js
Original file line number Diff line number Diff line change
Expand Up @@ -157,17 +157,46 @@ test("an agent taking a generated session id declares newSession", () => {
}
});

test("an idFrom pattern is a valid regular expression with one capture group", () => {
// A sample of each agent's real output, captured from the actual CLI. An
// idFrom pattern is only useful if it matches what the agent genuinely prints,
// so the fixture is the observed shape rather than an invented one.
const ID_SAMPLES = {
codex: ["session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"],
copilot: ["session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"],
cursor: ["chat_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"],
kimi: ["session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"],
droid: ['{"type":"result","result":"ok","session_id":"94c4334d-5e19-49e6-aea6-b45394e74370"}',
"94c4334d-5e19-49e6-aea6-b45394e74370"],
agy: ['{"conversation_id":"af730fe4-e36e-4146-a5c5-ba4fea33a325","status":"SUCCESS"}',
"af730fe4-e36e-4146-a5c5-ba4fea33a325"],
opencode: ['{"type":"text","part":{"sessionID":"ses_f5300fe7cffeZ4w7bSGLxUxGWS","text":"hi"}}',
"ses_f5300fe7cffeZ4w7bSGLxUxGWS"],
};

test("an idFrom pattern captures the session id from that agent's real output", () => {
for (const agent of BUILTIN_AGENTS) {
const src = agent.resume?.idFrom;
if (!src) continue;
const re = new RegExp(src, "i");
assert.equal(re.exec("session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233")?.[1],
"4f9a2c1b-33de-4a10-9f0e-7788aa112233",
const sample = ID_SAMPLES[agent.name];
assert.ok(sample, `${agent.name}: add a real output sample to ID_SAMPLES`);
const [stdout, expected] = sample;
assert.equal(new RegExp(src, "i").exec(stdout)?.[1], expected,
`${agent.name}: idFrom must capture the session id in group 1`);
}
});

test("every resuming agent can name the session it resumes", () => {
// Either the id is generated here and passed in (newSession), or it is
// scraped back out of the agent's own output (idFrom). Without one of the
// two, {{sessionId}} resolves to empty and the reply silently starts a new
// conversation — or worse, resumes whatever ran last.
for (const agent of BUILTIN_AGENTS) {
if (!agent.resume?.supported) continue;
assert.ok(agent.newSession || agent.resume.idFrom,
`${agent.name}: resume needs newSession or resume.idFrom to know which session to resume`);
}
});

test("codex remains the default judge and the new agents carry the reviewer role", async () => {
const cfg = await loadConfig(path.join(dir, "..", ".."));
assert.equal(judgeAgent(cfg).name, "codex", "adding agents must not move the default judge");
Expand Down
9 changes: 7 additions & 2 deletions test/kimi.test.js
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@
// that session without the profile flag Kimi refuses next to --session.
import { test } from "node:test";
import assert from "node:assert/strict";
import { mkdtemp, mkdir, writeFile, readFile, readdir, rm, access } from "node:fs/promises";
import { mkdtemp, mkdir, writeFile, readFile, readdir, rm, access, realpath } from "node:fs/promises";
import { tmpdir } from "node:os";
import path from "node:path";
import { execFileSync, spawnSync } from "node:child_process";
Expand Down Expand Up @@ -130,7 +130,12 @@ test("a review is parsed from the JSON stream, its session captured, and its liv
assert.deepEqual(made.args.slice(2, 4), ["--output-format", "stream-json"]);
assert.ok(path.isAbsolute(made.profile.path), "the profile must be an absolute path: the process starts in the worktree");
assert.equal(made.profile.exists, true, "{{packageDir}} must point at the installed package");
assert.equal(path.relative(root, made.profile.path).split(path.sep).join("/"), "lib/agents/kimi-reviewer.md");
// Compare real paths on both sides. On macOS the temp root is /var/... while
// {{packageDir}} resolves through /private/var/..., so the same file reached
// two ways compares unequal and path.relative answers with a ../../.. chain.
assert.equal(
path.relative(await realpath(root), await realpath(made.profile.path)).split(path.sep).join("/"),
"lib/agents/kimi-reviewer.md");

// What the console saw while it ran: words and tool calls, no JSON envelope.
const live = chunks.join("");
Expand Down
7 changes: 6 additions & 1 deletion test/opencode.test.js
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,12 @@ test('OpenCode installed CLI discovers, selects, and saves the judge, delivering
assert.equal(call.cwd, await realpath(worktree));
assert.deepEqual(call.args, ['run', '--dir', worktree, '--agent', 'plan', '--format', 'json', '--', prompt]);
assert.equal(call.permission.edit, 'deny');
assert.equal(agent.resume.supported, false);
// OpenCode resumes its own review when replying, and names the session
// explicitly: --session with the id scraped from its own event stream, never
// "the latest session", which could be an unrelated conversation.
assert.equal(agent.resume.supported, true);
assert.ok(agent.resume.argv.includes('{{sessionId}}'));
assert.ok(!agent.resume.argv.includes('--continue'));
assert.equal((await invoke(judgeAgent(cfg, 'opencode'))).verdict, 'clean');
call = JSON.parse(await readFile(path.join(worktree, 'call.json')));
assert.ok(call.args.includes('build')); assert.ok(call.args.includes('--auto'));
Expand Down
Loading