-
Notifications
You must be signed in to change notification settings - Fork 0
Panel Mode
Design agreed 2026-07-15 — not yet built. Run all enabled harnesses in parallel on the same prompt, present anonymized candidate responses, let the user select one; the selection is both the conversation's next step and a preference label for the thesis.
v1 implemented 2026-07-16 (services/harness/panel.ts + GUI
PanelConsole in Mission Control). Deltas from the spec below, all
user-decided:
-
agent_selections.outcomegainsuser_authored(+user_text): the operator typed their own response instead of picking a candidate — all four rejected AND the gold answer supplied, the strongest label. -
agent_selections.stimulus_event_idlinks a turn to theeventsrow it responded to (defect alerts). -
conversation_builds(one-to-many): conversations are independent of builds and span them; build references are tracked per-turn inagent_messages.metaand annotated in the transcript on build change. -
Event auto-trigger: a broker watcher polls alerting
defect_detectionevents and fans them out to the long-running console panel (auto-created;context.console=true) as DATA turns. Newest alert supersedes older ones while a turn is busy/pending — no queue. Historical events are never replayed on boot. - GUI: the right-docked sidebar chat is removed; Mission Control's PanelConsole is the single agent surface. Sonar ping + browser notification when candidates await selection or a defect alert fires.
- Candidate one-shots get a 240 s kill timeout so one hung CLI can't wedge the turn.
The between-build comparison (swap harness per build) yields unpaired
observations confounded by build-to-build differences. Panel mode yields
paired comparisons — four harnesses answering the identical prompt
against identical printer/database state — so "which harness earns user
validation" becomes a statistics question (Bradley-Terry/Elo over selections)
instead of an anecdote. Complements the objective metrics already recorded
(tool calls, tokens, latency, agent_actions).
-
The broker owns the conversation state, not the harnesses. One
canonical transcript per panel conversation (user turns + selected
answers). Every turn, the full canonical transcript is re-sent to each
enabled harness as a fresh headless one-shot — no native
--resumein panel mode. Harnesses are stateless response generators. (This also sidesteps the agy-ccross-talk bug.) v1 renders the transcript with a simple role-tagged text format; a syntax closer to the harnesses' native formats is a later refinement. - Selection = exactly one response. Only the chosen response enters the canonical transcript; the others are archived as candidates. If two harnesses give equivalent answers the user still picks one. No parallel execution ever: panel mode is propose-only (enforced by the advisor prompt) — if the selected response recommends a printer action, execution happens separately through a single-harness chat (or manually).
- Blinded selection. Candidates render as anonymized panes (A/B/C/D), shuffled server-side each turn; the harness identity is revealed only after the selection is persisted. The shown order is recorded so position bias can be checked. "Tie" / "none are good" outcomes are recordable — forced choice corrupts the signal.
-
Harness↔model correlation is accepted and documented. Claude Code
runs Claude models, Codex OpenAI models, Antigravity Gemini models —
selections measure the stack. The
modelfield is recorded on every candidate response for future reference. OpenCode is the deliberate exception: it is the multi-model harness (hosted models and the planned domain-adapted fine-tune), and OpenCode×{models} remains the model-isolating axis of the thesis.
user message ──▶ broker appends to canonical transcript
├─▶ claude (fresh one-shot, canonical context) ──┐
├─▶ opencode (same) ├─ SSE per pane,
├─▶ codex (same) │ shuffled A–D
└─▶ agy (same) ┘
user selects one (or tie/none) ──▶ agent_selections row
──▶ chosen text appended to canonical transcript
──▶ harness identities revealed in GUI
next turn repeats with the updated canonical transcript
Toggles: each panel has an enabled_harnesses set — disable two or three to
save tokens; a panel of one degrades to today's single-chat behavior. Panel
mode is opt-in per conversation (4× token cost).
Principle: a candidate response is just a normal conversation — a
one-turn agent_conversations row linked to its parent panel. Everything
that already exists (message normalization, agent_sessions per-turn stats,
transcript archival, the exporter and its credential gate) then covers panel
data with no special cases. One new table carries the preference labels.
Schema changes:
-- Parent link (generic — also future-proof for watchdog-spawned sessions)
ALTER TABLE agent_conversations
ADD COLUMN parent_id BIGINT REFERENCES agent_conversations(id),
ADD COLUMN parent_turn INT;
-- Preference labels, one row per panel turn
CREATE TABLE agent_selections (
id BIGSERIAL PRIMARY KEY,
conversation_id BIGINT NOT NULL REFERENCES agent_conversations(id), -- panel
turn INT NOT NULL,
outcome TEXT NOT NULL, -- chosen | tie | none
chosen_id BIGINT REFERENCES agent_conversations(id), -- candidate
candidate_ids BIGINT[] NOT NULL, -- all candidates this turn
shown_order JSONB NOT NULL, -- slot (A–D) → candidate id
reason TEXT, -- evaluator rubric note
decided_at TIMESTAMPTZ NOT NULL DEFAULT now(),
UNIQUE (conversation_id, turn)
);Row roles:
| Row | Table | Notes |
|---|---|---|
| Panel (canonical) |
agent_conversations, role='panel', harness='panel'
|
its agent_messages = user turns + selected answers only; provenance lives in agent_selections, not duplicated in messages |
| Candidate |
agent_conversations, role='panel:candidate', parent_id→panel, parent_turn=N |
one turn; harness/model as today; full normalized messages incl. tool calls |
| Candidate stats |
agent_sessions (role='panel:candidate') |
tokens, tool calls, latency, exit code — unchanged columns |
| Selection | agent_selections |
Bradley-Terry fits directly from this table |
Raw transcripts: each candidate archives as its own dir in the dataset repo,
source/sessions/<stamp>-panel<PID>-turn<N>-<harness>/. The canonical
transcript is DB-derived (no CLI run of its own).
conversations.jsonl keeps its one-row-per-conversation schema — panels and
candidates appear as ordinary rows with two new nullable fields
(parent_id, parent_turn) and the new roles. Candidate rows are directly
comparable across harness/model exactly like every other row, which is the
dataset's stated purpose.
New sibling export + HF config: data/selections.jsonl —
{id, conversation_id, turn, outcome, chosen_id, reason, decided_at,
shown_order,
candidates: [{conversation_id, harness, model, slot,
tokens_in, tokens_out, tool_calls, latency_ms}]}
(candidates is denormalized at export time by joining
agent_conversations + agent_sessions so the arena analysis needs only
this one file.) sls-export-conversations gains the second query + output;
the credential-leak gate applies to both files. Dataset README gains the
selections config and the panel / panel:candidate roles.
Must land before arena data collection:
-
spawnerror handlers + per-turn timeouts in the broker (one hung CLI must show a timeout card, not brick the panel). - Permission envelope: presets were removed 2026-07-18 — all harnesses now get the same full MCP tool surface, though built-in-tool leash still differs per CLI (codex/agy run with bypass flags; Claude's built-ins sit behind headless permission prompts). Selections must not measure who had the longer leash.
- Model capture for opencode/agy (currently NULL unless
--modelpassed).
Strongly recommended: server-side logging of ALL MCP tool calls in the
plugin (not just *_set) — the only harness-independent way to compare tool
behavior, since opencode/agy transcripts are plain text.
- Primary: selection outcomes → Bradley-Terry/Elo with confidence intervals; report ties/none rates.
- Bias checks: position effects from shown-order; selection reasons rubric (author is the sole evaluator — define the rubric up front).
- Secondary: per-candidate latency, tokens, tool-call counts, and objective correctness on database-grounded questions.