Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 8 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,10 +15,18 @@ parallel copies under `docs/` or `scripts/notes/`. At cut time: rename

### Added

### Added

- Plan and counsel workers require substance in Findings (files/paths,
acceptance criteria, non-goals, risks, ordered steps). Four headings
with stub Findings salvage as `incomplete-report`, not an attachable
plan. Implement and review envelope completeness is unchanged.
- Consecutive same-tool transcript calls collapse into one row with a count
chip (`· ×N`). Settled lanes use a past-tense head (`Grepped ×3 · "corbits"`).
`spawn_agent` stays one row per dispatch; `manage_tasks` paints no row.
- Pending `run_shell` rows stream up to three live output lines from a
bounded 8 KiB feed, then a last-three preview and a non-zero `exit N` at
settle.

### Fixed

Expand All @@ -31,7 +39,6 @@ parallel copies under `docs/` or `scripts/notes/`. At cut time: rename
ask is dropped. Do not poll `list_agents`.



## [0.3.21] - 2026-09-11

### Added
Expand Down
72 changes: 69 additions & 3 deletions docs/TUI.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@ down the left edge (`userBubbleLines` in `src/tui/stream.ts`). Each bubble
keeps one empty bar row above and below its text so the operator's voice
stays easy to find while scrolling through denser assistant and tool rows —
the pad is part of the bubble itself, not an extra turn-boundary gap, and
assistant/tool rows are unchanged.
assistant rows are unchanged; tool rows are lanes (see Tool lanes below).

Parent live reasoning paints through the existing thinking row — never a
third mid-turn stream lane. While `inference.thinking.delta` arrives,
Expand All @@ -75,6 +75,63 @@ same one row per turn (`reasoning-fold`); `inference.text.delta` grows the
open assistant streaming row in place. Worker spawn_agent-row thinking is a
separate path and is unchanged by this preview.

## Tool lanes

Transcript tool rows group by tool, not by sentence. A call for a tool the
previous row already represents folds onto that row instead of opening a new
one (`src/tui/tool-rows.ts`); the row narrates the newest call's subject and,
while calls are in flight, carries a dim count chip (`· ×N`) after the
subject. `spawn_agent` never folds — each dispatch is its own live anchor —
and `manage_tasks` paints no transcript row at all.

When the lane settles (its last outstanding answer landed), the head rewrites
to a past-tense count over the latest subject — `Grepped ×3 · "corbits"` —
with the count taken over calls, never over payload items (a lane of three
greps says `×3` even if the payloads returned forty matches in total; nothing
here can substantiate a payload total). Each call's own answer stays behind
the expand arrow. For an edit lane the per-call `+n/-n path` addenda remain
in the expanded body, so folding never buries which file each call touched.
Single-call rows (count <= 1) render exactly as they did before lanes.

The past tense is a map keyed by raw tool name (`src/tui/tool-formatter.ts`:
`grep` → `Grepped`, `read_file` → `Read`, `write_file` → `Wrote`,
`edit_file` → `Edited`, `run_shell` → `Ran`, `list_dir` → `Listed`,
`search_files` → `Searched`, …); an unknown tool falls back to its display
name.

Because a lane's row identity moves to the newest call, a lane also carries
the call ids it absorbed (`memberIds`, newest appended). A
result resolves its lane when its call id is the row's own id **or** one of
its members — this is what pairs a resumed transcript's parallel batch
(call, call, result, result) correctly. An id matching nothing still answers
nothing: it is appended as its own row, never folded onto the newest
same-name lane.

Alt+C on a lane copies the most recent call's full output (see the Alt+C
bullet under Clipboard and mouse).

### Shell output lanes

A pending `run_shell` row shows the command head and elapsed clock as today,
plus up to three dim tail lines of the command's live output and — when the
feed window holds more than three lines — the same dim `⋯ +N lines` elision
marker as the settle preview, `N` counting lines within the live window (the
feed keeps only the most recent 8 KiB). The output
travels through a polled bounded feed (`src/session/shell-output-feed.ts`),
not a reactor event: the plugin appends chunks (capped at an 8 KiB tail per
call, emitting at most once per 100 ms plus a final flush at settle), and
the product host's sticky poll reads the feed snapshot every 200 ms and
repaints the pending row frame-coalesced. Nothing from the feed is
persisted. When the feed is not wired (tests, the demo shell), the row
renders exactly as before — silent degradation.

At settle the live tail is replaced by a preview of the full output: its
last three lines plus a dim `⋯ +N lines` elision marker when more were
produced (the marker carries the count; the "N lines" stat is not painted
on shell rows). A non-zero exit adds an `exit N` stat; a zero exit adds
none. The existing expand idiom (Alt+E / click / the row arrow) reveals the
full output, hiding the preview; the idiom and the arrow are unchanged.

The prompt box's border carries the metadata that would otherwise cost a
titlebar row: the model label sits right-aligned in the top rule as
`profile · model · effort` (empty segments omitted), and a
Expand Down Expand Up @@ -249,7 +306,10 @@ clocks.
`runtime-bridge` paints each `spawn_agent` call as a transcript stream row for
**spawn / final / fail anchors**. While the agents strip is sticky, sticky-poll
`syncAgentProgress` rewrites are gated off so the transcript is not a dual live
rail. Ordinary in-flight tool rows keep their own elapsed clock
rail. A `spawn_agent` call never folds into a tool lane: each dispatch keeps
its own transcript row for its whole lifetime, because the row is the live
progress anchor, not just a call record. Ordinary in-flight tool rows keep
their own elapsed clock
(`syncToolElapsed`) without the current-tool suffix.

### Unprompted fleet reports
Expand Down Expand Up @@ -725,7 +785,9 @@ running its own selection. Two chords cover remaining copy needs:
(`enterCopyMode`) that resolves through the system clipboard port
(`src/tui/system-clipboard.ts` — a native helper binary per
platform, `pbcopy`/`clip`/`wl-copy`/`xclip`/`xsel`, falling back to an OSC
52 escape sequence when no helper is available, e.g. over SSH).
52 escape sequence when no helper is available, e.g. over SSH). On a
coalesced tool lane the copy resolves to the most recent call's full
output; a single-call row copies its own output, exactly as before.

Arrow keys never scroll anything — inside the prompt they are caret motion
or, at the buffer's edges, prompt-history recall; inside an open overlay's
Expand Down Expand Up @@ -787,6 +849,10 @@ terminal. It cannot observe:

- **Real paint.** Tests assert on the shell's in-memory row/rect state, not
on what a terminal emulator actually draws to a screen buffer.
- **The live shell tail's wall clock.** The feed itself is pure and the
cadence is injectable, so the live tail is headless-testable by driving a
fake feed and a fake clock through the same sync path the sticky poll
uses; what the harness cannot see is real-time emission timing.
- **Modifier reporting.** Whether a real terminal can report Shift+Enter,
Alt+letter, or similar modifier combinations depends on the terminal
negotiating the kitty keyboard protocol (or an equivalent) with the actual
Expand Down
6 changes: 6 additions & 0 deletions src/agent/posix-tool-plugins.ts
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,7 @@ import {
import type { PermissionGate } from "../permission/gate.js";
import { createWorktreeRootsProvider } from "../permission/worktree-roots.js";
import type { CompactionArchive } from "../session/compaction-archive.js";
import type { ShellOutputFeedMap } from "../session/shell-output-feed.js";

export interface CorePosixToolPluginsArgs {
cwd: string;
Expand All @@ -47,6 +48,9 @@ export interface CorePosixToolPluginsArgs {
// Live getter for the background-shell registry (run_shell background:true).
// Omitted makes background runs fail closed in shell-guard.
getBackgroundShellRegistry?: () => BackgroundShellRegistry | undefined;
// Live getter for the per-call bounded shell-output feeds the transcript
// polls for a running command's live tail. Omitted leaves the tail unwired.
getShellOutputFeeds?: () => ShellOutputFeedMap | undefined;
/** Primary-only evidence archive; workers omit this getter. */
getEvidenceArchive?: () => CompactionArchive | undefined;
}
Expand Down Expand Up @@ -86,6 +90,7 @@ export function buildCorePosixToolPlugins(
getContextDir,
shellEnv,
getBackgroundShellRegistry,
getShellOutputFeeds,
getEvidenceArchive,
} = args;
// Pre-gate sandboxes honor yolo mode so outside-workspace path tools and shell
Expand Down Expand Up @@ -119,6 +124,7 @@ export function buildCorePosixToolPlugins(
...(getBackgroundShellRegistry !== undefined
? { getBackgroundShellRegistry }
: {}),
...(getShellOutputFeeds !== undefined ? { getShellOutputFeeds } : {}),
}),
...(getEvidenceArchive !== undefined
? [evidenceArchiveSearchPlugin(getEvidenceArchive)]
Expand Down
10 changes: 10 additions & 0 deletions src/agent/tools.ts
Original file line number Diff line number Diff line change
Expand Up @@ -90,6 +90,7 @@ import {
createBackgroundShellRegistry,
type BackgroundShellExit,
} from "../shell/background-shell.js";
import { createShellOutputFeedMap } from "../session/shell-output-feed.js";
import { createListDirTool } from "../util/list-dir.js";
import {
createExaMCPWebFetchTool,
Expand Down Expand Up @@ -289,6 +290,9 @@ export interface AgentToolset {
callbacks: MCPConnectCallbacks,
signal?: AbortSignal,
) => Promise<void>;
// Per-call bounded live-output tails of foreground shells, polled by the
// transcript for each pending run_shell row's live lines.
shellOutputFeed: ReturnType<typeof createShellOutputFeedMap>;
// Connect one newly persisted server through the same lifecycle as startup MCP.
connectMCPServer: (
config: MCPServerConfig,
Expand Down Expand Up @@ -360,6 +364,10 @@ export async function createAgentToolset(
: {}),
});
const shellCollect = createShellCollectTool(backgroundShells);
// Per-call bounded live-output tails of foreground shells, polled by the TUI
// for each pending run_shell row's live lines. Workers get a map too; nothing
// reads it unless a transcript polls it (silent degradation).
const shellOutputFeed = createShellOutputFeedMap();
const sessionBlobReader =
getBlobReader !== undefined
? createLazyBlobReader(getBlobReader)
Expand Down Expand Up @@ -447,6 +455,7 @@ export async function createAgentToolset(
...(getEvidenceArchive !== undefined ? { getEvidenceArchive } : {}),
...(shellEnv !== undefined ? { shellEnv } : {}),
getBackgroundShellRegistry: () => backgroundShells,
getShellOutputFeeds: () => shellOutputFeed,
}),
});

Expand Down Expand Up @@ -1179,6 +1188,7 @@ export async function createAgentToolset(
return {
dynamicRunner,
connectMCP,
shellOutputFeed,
connectMCPServer: publicConnectMCPServer,
disconnectMCPServer: publicDisconnectMCPServer,
hasMCPServer: (name) =>
Expand Down
92 changes: 92 additions & 0 deletions src/plugins/shell-guard-plugin.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -10,9 +10,12 @@ import { spawnSync, type ChildProcess } from "node:child_process";
import { randomUUID } from "node:crypto";

import { createBackgroundShellRegistry } from "../shell/background-shell.js";
import { createShellOutputFeed } from "../session/shell-output-feed.js";

import {
BoundedShellOutput,
MAX_SHELL_OUTPUT_BYTES,
SHELL_FEED_EMIT_MS,
advertiseShellGuardTimeout,
resolveShellTimeoutMs,
reapLiveChildren,
Expand All @@ -37,6 +40,35 @@ describe("runGuardedShell", () => {
expect(output).toContain("hello");
});

test("a rate-limited second write reaches the feed before the process exits", async () => {
const feed = createShellOutputFeed();
let finished = false;
const running = runGuardedShell(
{ command: "echo first; sleep 0.02; echo second; sleep 0.4" },
neverAbort(),
undefined,
undefined,
(text) => {
feed.append(text);
},
).then((result) => {
finished = true;
return result;
});
const deadline = Date.now() + SHELL_FEED_EMIT_MS + 80;
while (
!feed.snapshot().includes("second") &&
Date.now() < deadline &&
!finished
) {
await Bun.sleep(10);
}
expect(finished).toBe(false);
expect(feed.snapshot()).toContain("second");
const result = await running;
expect(result.exitCode).toBe(0);
});

test("omitted timeout does not arm a timer", async () => {
const start = Date.now();
const { exitCode, timedOut, output } = await runGuardedShell(
Expand Down Expand Up @@ -309,6 +341,66 @@ describe("background run_shell (shellGuardPlugin)", () => {
expect(registry.runningCount()).toBe(0);
registry.disposeAll("test done");
});

test("an unwired shell-output feed spawns fine and paints no tail", async () => {
const handler = defined(
shellGuardPlugin(process.cwd(), undefined, undefined, {}).middleware,
)(fallback);
const result = await handler(
{ id: "fg2", name: "run_shell", arguments: { command: "echo hi" } },
neverAbort(),
);
expect(result.isError).toBeUndefined();
expect(String(result.content)).toContain("hi");
});

test("a wired feed receives the output tail at cadence with a final flush", async () => {
const feed = createShellOutputFeed();
let emits = 0;
const handler = defined(
shellGuardPlugin(process.cwd(), undefined, undefined, {
getShellOutputFeeds: () => {
const wrapped = {
append: (text: string) => {
emits += 1;
feed.append(text);
},
snapshot: () => feed.snapshot(),
clear: () => feed.clear(),
};
return {
forCall: () => wrapped,
get: () => wrapped,
drop: () => undefined,
};
},
}).middleware,
)(fallback);
const result = await handler(
{
id: "fg3",
name: "run_shell",
arguments: {
// Fifteen lines ~10 ms apart: far more chunk arrivals than one
// cadence window per 100 ms can allow. Without the Date.now() gate
// in emitPendingOutput every arrival emits (~16 emissions) and this
// ceiling fails — the assertion is what pins the cadence.
command:
"i=1; while [ $i -le 15 ]; do echo line$i; sleep 0.01; i=$((i+1)); done",
},
},
neverAbort(),
);
expect(result.isError).toBeUndefined();
// The final flush lands the tail (including the last line) in the feed.
expect(feed.snapshot()).toContain("line1");
expect(feed.snapshot()).toContain("line15");
// At most one emit per 100 ms of wall time (~300 ms with the pwd probe
// trailer), plus the final flush. Still far below the ~16 arrivals, so a
// broken cadence gate cannot pass.
const elapsedMs = 350;
expect(emits).toBeLessThanOrEqual(Math.ceil(elapsedMs / 100) + 1);
});
});

describe("advertiseShellGuardTimeout", () => {
Expand Down
Loading
Loading