Skip to content

Panel Mode

Peter Pak edited this page Jul 26, 2026 · 4 revisions

Panel Mode (Harness Arena)

Design agreed 2026-07-15 — not yet built. Run all enabled harnesses in parallel on the same prompt, present anonymized candidate responses, let the user select one; the selection is both the conversation's next step and a preference label for the thesis.

Status

v1 implemented 2026-07-16 (services/harness/panel.ts + GUI PanelConsole in Mission Control). Deltas from the spec below, all user-decided:

  • agent_selections.outcome gains user_authored (+ user_text): the operator typed their own response instead of picking a candidate — all four rejected AND the gold answer supplied, the strongest label.
  • agent_selections.stimulus_event_id links a turn to the events row it responded to (defect alerts).
  • conversation_builds (one-to-many): conversations are independent of builds and span them; build references are tracked per-turn in agent_messages.meta and annotated in the transcript on build change.
  • Event auto-trigger: a broker watcher polls alerting defect_detection events and fans them out to the long-running console panel (auto-created; context.console=true) as DATA turns. Newest alert supersedes older ones while a turn is busy/pending — no queue. Historical events are never replayed on boot.
  • GUI: the right-docked sidebar chat is removed; Mission Control's PanelConsole is the single agent surface. Sonar ping + browser notification when candidates await selection or a defect alert fires.
  • Candidate one-shots get a 240 s kill timeout so one hung CLI can't wedge the turn.

Why

The between-build comparison (swap harness per build) yields unpaired observations confounded by build-to-build differences. Panel mode yields paired comparisons — four harnesses answering the identical prompt against identical printer/database state — so "which harness earns user validation" becomes a statistics question (Bradley-Terry/Elo over selections) instead of an anecdote. Complements the objective metrics already recorded (tool calls, tokens, latency, agent_actions).

Decisions

  1. The broker owns the conversation state, not the harnesses. One canonical transcript per panel conversation (user turns + selected answers). Every turn, the full canonical transcript is re-sent to each enabled harness as a fresh headless one-shot — no native --resume in panel mode. Harnesses are stateless response generators. (This also sidesteps the agy -c cross-talk bug.) v1 renders the transcript with a simple role-tagged text format; a syntax closer to the harnesses' native formats is a later refinement.
  2. Selection = exactly one response. Only the chosen response enters the canonical transcript; the others are archived as candidates. If two harnesses give equivalent answers the user still picks one. No parallel execution ever: panel mode is propose-only (enforced by the advisor prompt) — if the selected response recommends a printer action, execution happens separately through a single-harness chat (or manually).
  3. Blinded selection. Candidates render as anonymized panes (A/B/C/D), shuffled server-side each turn; the harness identity is revealed only after the selection is persisted. The shown order is recorded so position bias can be checked. "Tie" / "none are good" outcomes are recordable — forced choice corrupts the signal.
  4. Harness↔model correlation is accepted and documented. Claude Code runs Claude models, Codex OpenAI models, Antigravity Gemini models — selections measure the stack. The model field is recorded on every candidate response for future reference. OpenCode is the deliberate exception: it is the multi-model harness (hosted models and the planned domain-adapted fine-tune), and OpenCode×{models} remains the model-isolating axis of the thesis.

Flow

user message ──▶ broker appends to canonical transcript
                 ├─▶ claude   (fresh one-shot, canonical context) ──┐
                 ├─▶ opencode (same)                                ├─ SSE per pane,
                 ├─▶ codex    (same)                                │  shuffled A–D
                 └─▶ agy      (same)                                ┘
user selects one (or tie/none) ──▶ agent_selections row
                                ──▶ chosen text appended to canonical transcript
                                ──▶ harness identities revealed in GUI
next turn repeats with the updated canonical transcript

Toggles: each panel has an enabled_harnesses set — disable two or three to save tokens; a panel of one degrades to today's single-chat behavior. Panel mode is opt-in per conversation (4× token cost).

Data model (settled 2026-07-15)

Principle: a candidate response is just a normal conversation — a one-turn agent_conversations row linked to its parent panel. Everything that already exists (message normalization, agent_sessions per-turn stats, transcript archival, the exporter and its credential gate) then covers panel data with no special cases. One new table carries the preference labels.

Schema changes:

-- Parent link (generic — also future-proof for watchdog-spawned sessions)
ALTER TABLE agent_conversations
  ADD COLUMN parent_id   BIGINT REFERENCES agent_conversations(id),
  ADD COLUMN parent_turn INT;

-- Preference labels, one row per panel turn
CREATE TABLE agent_selections (
  id              BIGSERIAL PRIMARY KEY,
  conversation_id BIGINT NOT NULL REFERENCES agent_conversations(id), -- panel
  turn            INT NOT NULL,
  outcome         TEXT NOT NULL,      -- chosen | tie | none
  chosen_id       BIGINT REFERENCES agent_conversations(id), -- candidate
  candidate_ids   BIGINT[] NOT NULL,  -- all candidates this turn
  shown_order     JSONB NOT NULL,     -- slot (A–D) → candidate id
  reason          TEXT,               -- evaluator rubric note
  decided_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
  UNIQUE (conversation_id, turn)
);

Row roles:

Row Table Notes
Panel (canonical) agent_conversations, role='panel', harness='panel' its agent_messages = user turns + selected answers only; provenance lives in agent_selections, not duplicated in messages
Candidate agent_conversations, role='panel:candidate', parent_id→panel, parent_turn=N one turn; harness/model as today; full normalized messages incl. tool calls
Candidate stats agent_sessions (role='panel:candidate') tokens, tool calls, latency, exit code — unchanged columns
Selection agent_selections Bradley-Terry fits directly from this table

Raw transcripts: each candidate archives as its own dir in the dataset repo, source/sessions/<stamp>-panel<PID>-turn<N>-<harness>/. The canonical transcript is DB-derived (no CLI run of its own).

Dataset integration (Agentic-SLS-Conversations)

conversations.jsonl keeps its one-row-per-conversation schema — panels and candidates appear as ordinary rows with two new nullable fields (parent_id, parent_turn) and the new roles. Candidate rows are directly comparable across harness/model exactly like every other row, which is the dataset's stated purpose.

New sibling export + HF config: data/selections.jsonl —

{id, conversation_id, turn, outcome, chosen_id, reason, decided_at,
 shown_order,
 candidates: [{conversation_id, harness, model, slot,
               tokens_in, tokens_out, tool_calls, latency_ms}]}

(candidates is denormalized at export time by joining agent_conversations + agent_sessions so the arena analysis needs only this one file.) sls-export-conversations gains the second query + output; the credential-leak gate applies to both files. Dataset README gains the selections config and the panel / panel:candidate roles.

Prerequisites (from the harness review, 2026-07-15)

Must land before arena data collection:

  • spawn error handlers + per-turn timeouts in the broker (one hung CLI must show a timeout card, not brick the panel).
  • Permission envelope: presets were removed 2026-07-18 — all harnesses now get the same full MCP tool surface, though built-in-tool leash still differs per CLI (codex/agy run with bypass flags; Claude's built-ins sit behind headless permission prompts). Selections must not measure who had the longer leash.
  • Model capture for opencode/agy (currently NULL unless --model passed).

Strongly recommended: server-side logging of ALL MCP tool calls in the plugin (not just *_set) — the only harness-independent way to compare tool behavior, since opencode/agy transcripts are plain text.

Evaluation plan

  • Primary: selection outcomes → Bradley-Terry/Elo with confidence intervals; report ties/none rates.
  • Bias checks: position effects from shown-order; selection reasons rubric (author is the sole evaluator — define the rubric up front).
  • Secondary: per-candidate latency, tokens, tool-call counts, and objective correctness on database-grounded questions.

Clone this wiki locally