diff --git a/CHANGELOG.md b/CHANGELOG.md index 4bb96f9..60987df 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,14 @@ results and metadata, and rejects non-JSON or empty output; the console shows the decoded words while the agent runs. `{{packageDir}}` in a command names the installed package directory, for files that ship with it. +- Claude, Droid, Antigravity and OpenCode now resume their own review conversation + when replying to verdicts, instead of starting fresh with their findings quoted + back. Eight of eleven agents now resume. Each names an explicit session id — + Claude via an assigned `--session-id`, the others read from structured output — + so a reply can never land in an unrelated conversation. +- Droid and Antigravity are read through their JSON output modes, which is where + each reports its session id. Antigravity's reported status is now checked, so a + run that ends early cannot be read as a sign-off. ## 0.6.0 — 2026-09-16 diff --git a/docs/configuration.md b/docs/configuration.md index 3b4ee43..25b6b0b 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -115,6 +115,34 @@ your environment. Run records include prompts and outputs, so redact them before ### Authentication, versions, and conversations +A reviewer's reply is a turn in the conversation that raised the findings, so +the agent can see what it said. Every resume names an explicit session id — +never "the latest session", which would answer whatever ran most recently in +that worktree rather than this review. + +An id reaches Jury one of two ways. Some CLIs accept one we generate +(`--session-id`), so the conversation is identified before it exists. The rest +print one in structured output, which is read back with `resume.idFrom`. An id +is never inferred from assistant prose: a model that happens to mention a UUID +is not reporting its session, and resuming on that would deliver a verdict into +an unrelated conversation. + +| Agent | Session id | Resumes | +|---|---|---| +| Claude | assigned `--session-id` | yes | +| Grok | assigned `--session-id` | yes | +| Qwen | assigned `--session-id` | yes | +| Codex | printed by `exec` | yes | +| Droid | `session_id` in `-o json` | yes | +| Antigravity | `conversation_id` in `--output-format json` | yes | +| OpenCode | `sessionID` in the JSON event stream | yes | +| Kimi Code | printed at the end of the `stream-json` output | yes | +| Amp, Cursor, Copilot | — | no; replies start fresh with the finding context | + +Agents that do not resume lose nothing in substance: `buildReply` quotes their +own prior findings back to them. The judge never resumes at all — each finding +is triaged in its own session so one verdict cannot anchor the next. + - Kimi Code: verified CLI 0.39.1 and 0.42.0 (`npm install -g @moonshot-ai/kimi-code`); run `kimi login`, or put an API key in `~/.kimi-code/config.toml`, whose `default_model` is the model used. Prompt mode (`kimi -p`) rejects `--plan`, `--yolo` and `--auto` and approves every tool call itself, so the reviewer boundary is the shipped agent profile: it wraps Kimi's default prompt, removes the Edit and Write tools, and allows only the read-only `explore` sub-agent (the default `coder` sub-agent can write). The shell stays available, so the profile is not a sandbox. The report is read from `--output-format stream-json`, and replies resume the exact session id Kimi prints at the end of the stream; the profile flag is omitted on resume because Kimi refuses it next to `--session`, and the session keeps the agent it was created with. Kimi keeps its own session history per working directory under `~/.kimi-code`. The legacy Python `kimi-cli` installs an executable of the same name but takes different flags (`kimi --version` prints 1.x for it and 0.x for Kimi Code); to keep using it, override the entry in `jury.config.json`: ```json @@ -123,6 +151,6 @@ your environment. Run records include prompts and outputs, so redact them before - Cursor: verified `cursor-agent` 2025.10.28-0a91dc2; run `cursor-agent login`. The generic executable `agent` may belong to another product, so Jury uses `cursor-agent`. Final JSON must explicitly report success. Replies start fresh. - Copilot: verified CLI 0.0.392 flags; authenticate with interactive `/login`. A reviewer can inspect files and Git but cannot use write tools; the judge allows tools. Local permission configuration remains trusted. Replies start fresh. - Qwen: targets CLI 0.23.4 (`npm install -g @qwen-code/qwen-code`); run `qwen` and complete `/auth` before headless use. It runs in the process worktree, not merely an added access directory. Reviews use assigned UUIDs, and replies resume only that UUID; no implicit latest session. Final JSON must explicitly report success. -- Amp: verified CLI 0.0.1788739286; run `amp login`. Threads are private and IDE context is disabled. Both roles can execute tools automatically; replies start fresh with their own finding context. +- Amp: verified CLI 0.0.1788739286; run `amp login`. Threads are private and IDE context is disabled. Both roles can execute tools automatically; replies start fresh with their own finding context. (`--stream-json` does expose a thread id, so resume is possible once that output mode is parsed.) A missing executable, nonzero exit, timeout, or unsuccessful structured result never approves a review. Git command allowlists are CLI tool permissions, not OS sandboxes; use trusted local agent configuration. Model credentials and service availability are prerequisites for live inference. diff --git a/lib/agent-schema.js b/lib/agent-schema.js index 9c5fa9b..114d43a 100644 --- a/lib/agent-schema.js +++ b/lib/agent-schema.js @@ -14,7 +14,7 @@ export const PROMPT_DELIVERY = ["argv", "file", "stdin"]; export const CWD_MODES = ["worktree", "flag"]; /** How much of stdout is the agent's report. */ -export const REPORT_MODES = ["whole", "tail", "opencode-json", "result-json", "kimi-json"]; +export const REPORT_MODES = ["whole", "tail", "opencode-json", "result-json", "kimi-json", "agy-json"]; /** * How strongly the reviewer invocation is prevented from writing. diff --git a/lib/agents.js b/lib/agents.js index 6d48796..693bc46 100644 --- a/lib/agents.js +++ b/lib/agents.js @@ -175,6 +175,20 @@ export function extractReport(stdout, mode = "whole") { } return result.result.trim(); } + if (mode === "agy-json") { + let data; + try { data = JSON.parse(text); } catch { throw new Error("Invalid Antigravity JSON output"); } + // Antigravity reports its own outcome. A timeout that still exits 0 with a + // partial answer arrives as a non-SUCCESS status, and reading the prose + // regardless is how a truncated review gets accepted as a sign-off. + if (data?.status !== "SUCCESS") { + throw new Error(`Antigravity did not finish: ${data?.status ?? "missing status"}`); + } + if (typeof data.response !== "string" || !data.response.trim()) { + throw new Error("Antigravity returned no response text"); + } + return data.response.trim(); + } if (mode === "opencode-json") { let parts = []; let messageId; diff --git a/lib/agents/agy.json b/lib/agents/agy.json index 77b044f..d3d3e5f 100644 --- a/lib/agents/agy.json +++ b/lib/agents/agy.json @@ -5,13 +5,38 @@ "promptDelivery": "argv", "cwd": "worktree", "argv": [ - "agy", "--dangerously-skip-permissions", "--add-dir", "{{worktree}}", - "--print-timeout", "20m", "--print", "{{promptText}}" + "agy", + "--dangerously-skip-permissions", + "--add-dir", + "{{worktree}}", + "--print-timeout", + "20m", + "--output-format", + "json", + "--print", + "{{promptText}}" ], "sandbox": "none", "sandboxNote": "--dangerously-skip-permissions approves every tool, including writes", - "resume": { "supported": false, "reason": "print mode starts a fresh session per run" }, - "report": "whole", + "resume": { + "supported": true, + "argv": [ + "agy", + "--dangerously-skip-permissions", + "--add-dir", + "{{worktree}}", + "--print-timeout", + "20m", + "--output-format", + "json", + "--conversation", + "{{sessionId}}", + "--print", + "{{promptText}}" + ], + "idFrom": "\"conversation_id\"\\s*:\\s*\"([^\"]+)\"" + }, + "report": "agy-json", "expectSeconds": 420, "install": "npm install -g @google/antigravity-cli", "docs": "https://antigravity.google/docs/cli" diff --git a/lib/agents/claude.json b/lib/agents/claude.json index 51f9196..27b8595 100644 --- a/lib/agents/claude.json +++ b/lib/agents/claude.json @@ -5,21 +5,48 @@ "promptDelivery": "argv", "cwd": "worktree", "argv": [ - "claude", "-p", "{{promptText}}", - "--permission-mode", "plan", - "--disallowedTools", "Edit,Write,NotebookEdit", - "--add-dir", "{{worktree}}" + "claude", + "-p", + "{{promptText}}", + "--session-id", + "{{sessionId}}", + "--permission-mode", + "plan", + "--disallowedTools", + "Edit,Write,NotebookEdit", + "--add-dir", + "{{worktree}}" ], "judgeArgv": [ - "claude", "-p", "{{promptText}}", - "--permission-mode", "acceptEdits", - "--add-dir", "{{worktree}}" + "claude", + "-p", + "{{promptText}}", + "--permission-mode", + "acceptEdits", + "--add-dir", + "{{worktree}}" ], "sandbox": "plan", "sandboxNote": "reviews run in plan mode with Edit, Write and NotebookEdit withheld", - "resume": { "supported": false, "reason": "each finding is judged on its own merits" }, + "resume": { + "supported": true, + "argv": [ + "claude", + "-p", + "{{promptText}}", + "--resume", + "{{sessionId}}", + "--permission-mode", + "plan", + "--disallowedTools", + "Edit,Write,NotebookEdit", + "--add-dir", + "{{worktree}}" + ] + }, "report": "whole", "expectSeconds": 900, "install": "npm install -g @anthropic-ai/claude-code", - "docs": "https://docs.claude.com/en/docs/claude-code/cli-reference" + "docs": "https://docs.claude.com/en/docs/claude-code/cli-reference", + "newSession": true } diff --git a/lib/agents/droid.json b/lib/agents/droid.json index feec4f8..6a11bad 100644 --- a/lib/agents/droid.json +++ b/lib/agents/droid.json @@ -4,11 +4,39 @@ "role": "reviewer", "promptDelivery": "file", "cwd": "flag", - "argv": ["droid", "exec", "--cwd", "{{worktree}}", "--auto", "medium", "-f", "{{promptFile}}"], + "argv": [ + "droid", + "exec", + "--cwd", + "{{worktree}}", + "--auto", + "medium", + "-o", + "json", + "-f", + "{{promptFile}}" + ], "sandbox": "plan", "sandboxNote": "--auto medium withholds destructive actions but permits edits", - "resume": { "supported": false, "reason": "exec output carries no session id" }, - "report": "whole", + "resume": { + "supported": true, + "argv": [ + "droid", + "exec", + "--cwd", + "{{worktree}}", + "--auto", + "medium", + "-o", + "json", + "-s", + "{{sessionId}}", + "-f", + "{{promptFile}}" + ], + "idFrom": "\"session_id\"\\s*:\\s*\"([^\"]+)\"" + }, + "report": "result-json", "expectSeconds": 130, "install": "curl -fsSL https://app.factory.ai/cli | sh", "docs": "https://docs.factory.ai/cli/getting-started/quickstart" diff --git a/lib/agents/opencode.json b/lib/agents/opencode.json index 6263169..925ab6b 100644 --- a/lib/agents/opencode.json +++ b/lib/agents/opencode.json @@ -38,8 +38,22 @@ "sandbox": "tools", "sandboxNote": "OPENCODE_PERMISSION denies edit, task and all bash but a read-only git allowlist", "resume": { - "supported": false, - "reason": "fresh conversation with finding context; never resume the latest session" + "supported": true, + "argv": [ + "opencode", + "run", + "--dir", + "{{worktree}}", + "--agent", + "plan", + "--format", + "json", + "--session", + "{{sessionId}}", + "--", + "{{promptText}}" + ], + "idFrom": "\"sessionID\"\\s*:\\s*\"([^\"]+)\"" }, "report": "opencode-json", "expectSeconds": 600, diff --git a/test/agy.test.js b/test/agy.test.js index 68aa981..a706d28 100644 --- a/test/agy.test.js +++ b/test/agy.test.js @@ -17,7 +17,9 @@ test('Antigravity is discoverable without configuration and receives literal pro const bin = path.join(dir, 'bin'); await mkdir(worktree); await mkdir(bin); const executable = path.join(bin, 'agy'); - await writeFile(executable, `#!${process.execPath}\nimport('node:fs').then(fs => { fs.writeFileSync('invocation.json', JSON.stringify({cwd:process.cwd(), args:process.argv.slice(2)})); console.log('NO NEW FINDINGS'); });\n`, { mode: 0o755 }); + // Antigravity is read through --output-format json, so the stub answers in + // that shape: a bare line of prose would no longer be a valid report. + await writeFile(executable, `#!${process.execPath}\nimport('node:fs').then(fs => { fs.writeFileSync('invocation.json', JSON.stringify({cwd:process.cwd(), args:process.argv.slice(2)})); console.log(JSON.stringify({conversation_id:'af730fe4-e36e-4146-a5c5-ba4fea33a325', status:'SUCCESS', response:'NO NEW FINDINGS'})); });\n`, { mode: 0o755 }); const cfg = await loadConfig(worktree, { globalFile: path.join(dir, 'missing.json') }); const agent = reviewers(cfg).find(a => a.name === 'agy'); assert.ok(agent); @@ -29,7 +31,10 @@ test('Antigravity is discoverable without configuration and receives literal pro assert.equal(result.verdict, 'clean'); const call = JSON.parse(await readFile(path.join(worktree, 'invocation.json'))); assert.equal(call.cwd, await realpath(worktree)); - assert.deepEqual(call.args, ['--dangerously-skip-permissions', '--add-dir', worktree, '--print-timeout', '20m', '--print', prompt]); + assert.deepEqual(call.args, ['--dangerously-skip-permissions', '--add-dir', worktree, '--print-timeout', '20m', '--output-format', 'json', '--print', prompt]); + // The conversation id is captured so the reply resumes this exact review + // rather than whatever conversation happens to be most recent. + assert.equal(result.sessionId, 'af730fe4-e36e-4146-a5c5-ba4fea33a325'); const env = { ...process.env, HOME: dir, USERPROFILE: dir, PATH: `${bin}${path.delimiter}${process.env.PATH}` }; const listing = spawnSync(process.execPath, [cli, 'agents'], { cwd: worktree, env, encoding: 'utf8' }).stdout; assert.match(listing, /ok\s+agy\s+reviewer/); diff --git a/test/builtin-agents.test.js b/test/builtin-agents.test.js index cd0e789..2301699 100644 --- a/test/builtin-agents.test.js +++ b/test/builtin-agents.test.js @@ -157,17 +157,46 @@ test("an agent taking a generated session id declares newSession", () => { } }); -test("an idFrom pattern is a valid regular expression with one capture group", () => { +// A sample of each agent's real output, captured from the actual CLI. An +// idFrom pattern is only useful if it matches what the agent genuinely prints, +// so the fixture is the observed shape rather than an invented one. +const ID_SAMPLES = { + codex: ["session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"], + copilot: ["session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"], + cursor: ["chat_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"], + kimi: ["session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"], + droid: ['{"type":"result","result":"ok","session_id":"94c4334d-5e19-49e6-aea6-b45394e74370"}', + "94c4334d-5e19-49e6-aea6-b45394e74370"], + agy: ['{"conversation_id":"af730fe4-e36e-4146-a5c5-ba4fea33a325","status":"SUCCESS"}', + "af730fe4-e36e-4146-a5c5-ba4fea33a325"], + opencode: ['{"type":"text","part":{"sessionID":"ses_f5300fe7cffeZ4w7bSGLxUxGWS","text":"hi"}}', + "ses_f5300fe7cffeZ4w7bSGLxUxGWS"], +}; + +test("an idFrom pattern captures the session id from that agent's real output", () => { for (const agent of BUILTIN_AGENTS) { const src = agent.resume?.idFrom; if (!src) continue; - const re = new RegExp(src, "i"); - assert.equal(re.exec("session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233")?.[1], - "4f9a2c1b-33de-4a10-9f0e-7788aa112233", + const sample = ID_SAMPLES[agent.name]; + assert.ok(sample, `${agent.name}: add a real output sample to ID_SAMPLES`); + const [stdout, expected] = sample; + assert.equal(new RegExp(src, "i").exec(stdout)?.[1], expected, `${agent.name}: idFrom must capture the session id in group 1`); } }); +test("every resuming agent can name the session it resumes", () => { + // Either the id is generated here and passed in (newSession), or it is + // scraped back out of the agent's own output (idFrom). Without one of the + // two, {{sessionId}} resolves to empty and the reply silently starts a new + // conversation — or worse, resumes whatever ran last. + for (const agent of BUILTIN_AGENTS) { + if (!agent.resume?.supported) continue; + assert.ok(agent.newSession || agent.resume.idFrom, + `${agent.name}: resume needs newSession or resume.idFrom to know which session to resume`); + } +}); + test("codex remains the default judge and the new agents carry the reviewer role", async () => { const cfg = await loadConfig(path.join(dir, "..", "..")); assert.equal(judgeAgent(cfg).name, "codex", "adding agents must not move the default judge"); diff --git a/test/kimi.test.js b/test/kimi.test.js index 07b1dc0..7783ad1 100644 --- a/test/kimi.test.js +++ b/test/kimi.test.js @@ -8,7 +8,7 @@ // that session without the profile flag Kimi refuses next to --session. import { test } from "node:test"; import assert from "node:assert/strict"; -import { mkdtemp, mkdir, writeFile, readFile, readdir, rm, access } from "node:fs/promises"; +import { mkdtemp, mkdir, writeFile, readFile, readdir, rm, access, realpath } from "node:fs/promises"; import { tmpdir } from "node:os"; import path from "node:path"; import { execFileSync, spawnSync } from "node:child_process"; @@ -130,7 +130,12 @@ test("a review is parsed from the JSON stream, its session captured, and its liv assert.deepEqual(made.args.slice(2, 4), ["--output-format", "stream-json"]); assert.ok(path.isAbsolute(made.profile.path), "the profile must be an absolute path: the process starts in the worktree"); assert.equal(made.profile.exists, true, "{{packageDir}} must point at the installed package"); - assert.equal(path.relative(root, made.profile.path).split(path.sep).join("/"), "lib/agents/kimi-reviewer.md"); + // Compare real paths on both sides. On macOS the temp root is /var/... while + // {{packageDir}} resolves through /private/var/..., so the same file reached + // two ways compares unequal and path.relative answers with a ../../.. chain. + assert.equal( + path.relative(await realpath(root), await realpath(made.profile.path)).split(path.sep).join("/"), + "lib/agents/kimi-reviewer.md"); // What the console saw while it ran: words and tool calls, no JSON envelope. const live = chunks.join(""); diff --git a/test/opencode.test.js b/test/opencode.test.js index 3ad0a4c..a72d815 100644 --- a/test/opencode.test.js +++ b/test/opencode.test.js @@ -60,7 +60,12 @@ test('OpenCode installed CLI discovers, selects, and saves the judge, delivering assert.equal(call.cwd, await realpath(worktree)); assert.deepEqual(call.args, ['run', '--dir', worktree, '--agent', 'plan', '--format', 'json', '--', prompt]); assert.equal(call.permission.edit, 'deny'); - assert.equal(agent.resume.supported, false); + // OpenCode resumes its own review when replying, and names the session + // explicitly: --session with the id scraped from its own event stream, never + // "the latest session", which could be an unrelated conversation. + assert.equal(agent.resume.supported, true); + assert.ok(agent.resume.argv.includes('{{sessionId}}')); + assert.ok(!agent.resume.argv.includes('--continue')); assert.equal((await invoke(judgeAgent(cfg, 'opencode'))).verdict, 'clean'); call = JSON.parse(await readFile(path.join(worktree, 'call.json'))); assert.ok(call.args.includes('build')); assert.ok(call.args.includes('--auto'));