feat: Phase 0 supervisor, e2e harness, and radius/v1 protocol versioning - #1
Merged
Merged
Conversation
Finishes the two Phase 0 items that were buildable without a product
decision, and records what the resulting evidence does and does not prove.
AgentSupervisor (src/supervisor.ts, 7 tests)
Three properties, each a way unattended agents go wrong:
- never two writers on one stream; start() refuses while running
- nothing observed left unpersisted; stop() drains even when dispose throws
- no process residency — a wedged harness raises StopTimeoutError rather
than hanging the host forever
The residency test spawns a REAL OS process, records its pid, stops, then
polls until the pid is gone. A mock asserting "dispose was called" would
pass while the process lived on, which is the bug that property exists for.
kill -9 e2e (test/e2e-kill.test.ts), run against a real Claude session
10 records durable after a real SIGKILL: contiguous from seq 0, all three
oar record kinds, replayed identically by a fresh store. The prompt, the
accepted response, system/init, the model's assistant frame and a terminal
result/success all survived — stamped with OUR stream name, not oar's native
id, so the untrusted-name property holds in real conditions.
Stated plainly: on this machine the turn COMPLETED before the kill, even at
a 9s delay. So this proves records survive a hard kill; it does NOT yet
prove a kill *during streaming* is safe. The run now reports which case it
hit, and RADIUS_E2E_KILL_MS exists to aim at a colder one. Mid-generation
remains unproven.
Also measured, because D-004 was blocking Phase 0's own gate:
cargo build --lib succeeds (~11s); cargo build --all-targets FAILS. Five of
six declared Cargo targets are missing — the whole examples/ directory was
never vendored. So a path dependency works and a built binary does not.
check:notices now parses the manifest and fails on an unrecorded missing
target, so a re-vendor cannot silently drop one.
Gates: 8/8 green. 39 tests.
Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
…nforced Phase 1.1 and 1.4: the two pieces whose logic can be proved exhaustively without a platform runtime. packages/protocol, 24 tests. Principal (1.1) The core type, with the deliberate omission made STRUCTURAL rather than conventional: no token/secret/apiKey field, and parsePrincipal rejects unknown fields so a credential cannot ride in through an untyped JSON bag. DATA-MODEL.md §4 asks for "a test that scans for it. Not a code-review convention" — so the tests grep serialized principals AND grants for nine secret shapes, with a positive control that plants a real-looking AWS key and requires detection. A scanner matching nothing would otherwise pass every other test and look green. CapabilityGrant.expiresAt is required in the type, so a permanent grant is unrepresentable rather than merely discouraged (DATA-MODEL §2, D-008). Lease manager (1.4) The primitive oar explicitly declines. Exclusive claims, epoch fencing, and expiry computed from process.hrtime — never Date.now(), which the user owns and can rewind with one `date` command. The restart case is the subtle one and D-008 names it: hrtime resets on restart, so restore() adopts a persisted anchor and DROPS any lease whose window elapsed while the process was down. Reviving it would let a rewound clock resurrect a dead lease. What this changes about the project's honesty: check:invariants now reports "11 specified only, 3 enforced in code, 1 verified against vendored source" instead of implying all 15 were prose. I3 (expired lease refused), I5 (stale epoch rejected) and I14 (clock rollback cannot extend a lease) moved from spec to enforced code. Each is mutation-verified: removing the epoch check, dropping the expiry check, switching expiry to Date.now(), freezing the epoch counter, and making the parser permissive each fail the suite. Two real bugs were found by these tests, not by review: the monotonic clock's continuity formula was wrong across a restart (time jumped to 0, which would have expired every live lease on any host bounce), and LeaseManager had no way to rehydrate leases at all — a Durable Object could not have restarted. Gates: 8/8 green. 63 tests. Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
Phase 1.2, minus the part that cannot honestly be built yet.
WHAT IS BUILT AND TESTED (28 tests)
TokenStore tokens scoped to (principal, launchId, capability); 256-bit
CSPRNG; only sha256(token) is retained, so a leaked table
yields nothing replayable (DATA-MODEL.md §4). Monotonic
expiry, launch binding, revocation without agent restart,
revokeLaunch as a lost-laptop kill switch, and narrowing of a
live token that can only ever shrink.
Policy deny by default. An ungranted capability is refused, and
every decision is audited INCLUDING denials — a log that
records only successes shows you nothing when something goes
wrong. D-014 enforced here as well as in the schema, so a
caller cannot skip the table. An expired grant reports as
expired rather than as a generic refusal, because an operator
debugging needs the real reason.
Provider scoped behind a narrow interface. ResolvedCredential is
opaque — no accessor, no toString — so a raw key is one
careless JSON.stringify away from a log line otherwise, and
the only way to use one is to hand it back to the adapter that
issued it.
WHAT IS NOT BUILT, AND WHY
Real credential scoping against Anthropic/OpenAI/Google/xAI. That needs
live provider keys. A broker that looks finished while credentials escape
is the worst thing this project could ship, so the boundary is isolated and
every provider path is marked unverified rather than stubbed into looking
real. Unscopable providers are REFUSED (assertScopable) rather than given a
raw key — that check is what stops the product quietly regressing to YOLO.
Still to prove with a live key: that a scoped credential cannot act beyond
its scope; that revocation stops the upstream call and not just our
bookkeeping; and that the credential is unreachable from agent context, a
heap dump, and error messages.
The loopback HTTP proxy is not built either — it is plumbing around a
decision engine that is now correct, and building it before the provider
boundary is settled would put an unverified surface in front of a real key.
MUTATION-VERIFIED (7/7 caught)
storing the raw token, dropping the monotonic expiry check, making narrow
widen, allowing unknown capabilities, letting an agent approve itself,
changing the denial reason, and neutering revokeLaunch.
Gates: 8/8 green. 91 tests.
Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
The remaining locally-buildable work. 127 tests, 8/8 gates. Sandbox policy (1.3) — deny by default, and honest about measurement The design problem I8 poses is "measured, or not claimed", and the way most code fails it is by behaving safely while its docs say "sandboxed". So `unknown` and `unmeasured` are first-class values, claimsIsolation() returns false for both, and an unprobed runtime is REFUSED at decision time — not merely described cautiously. A module that reads well and reports a sandbox nobody probed is exactly how a claim becomes fiction. Unprobed > explicit grant > escalate > deny, and an escalation carries the exact request so the required approvals row cannot be lost in translation. Approvals (D-013) — exactly-once, by construction One slot keyed by requestId, so a client retry physically cannot produce a second row. A replay carrying DIFFERENT content is refused rather than overwritten, because overwriting would let an agent re-approve a decision it was never granted. This is the only exactly-once claim in the system, and PROTOCOL.md is emphatic it must not be blurred into a general one. Cost ledger (Phase 4 gate) — a cap that holds under fault injection The gate says "cannot be exceeded, including under retry and partial failure". Retries cost twice, partial failure splits reservation from settlement, and concurrent turns race — so the cap is enforced at RESERVATION time against committed+reserved, and reserving the upper bound means a partial result cannot overshoot. A repeated settlement is a no-op, a crash between reserve and send releases headroom rather than leaking it, and a cost above the reservation is refused rather than absorbed. A real bug the tests caught: reservedFor counted SETTLED reservations, so headroom was charged twice — available() understated the budget and would have silently throttled an agent that still had money. MUTATION-VERIFIED (7/7 caught) a retry creating a second row, an agent approving itself, a cap ignoring in-flight reservations, a double settlement, an unprobed runtime being allowed, claimsIsolation always returning true, and describeIsolation claiming a sandbox for an unprobed runtime. STILL UNVERIFIED, and it stays that way: real provider credential scoping, and the actual isolation strength of any runtime. The policy is built and tested; the boundary it guards is not yet demonstrated with a live key or a real sandbox probe. Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
Closes the remaining locally-buildable work. 148 tests, 8/8 gates. Durability harness (milestone 0.4) Truncates the durable stream file at EVERY byte offset and asserts what survives is always a contiguous prefix [0..k] — no gap, no duplicate, no half-written record — then replays a truncated stream and asserts the resume delivers every later record exactly once. It does not assert "nothing is lost", because that is false and RecordWriter documents why: a record observed but not yet fsync'd dies with the process. A corrupt prefix is unacceptable; a lost tail is expected. Asserting the stronger false thing would have made this harness look better and be worse. Writing it caught the harness's own floating promises first — `void store.append(...)` raced the teardown. The exact hazard RecordWriter exists to prevent, caught by the thing built to detect it. Phase 2 logic (sync.ts) Tenant resolution, outbox-then-settle, and cursor merge — the parts of the control plane that are logic rather than deployment. Tenant resolution refuses an unknown token rather than defaulting, because a default tenant is a silent cross-tenant read. The refusal names why it happens at the edge: resolving inside a DO means the DO was already addressed by an id the client chose. Outbox settle is idempotent in both directions — a double settle is a no-op, and re-enqueuing an id does not clobber a row another worker is still settling. The property is redelivery, never loss. CursorMerger is explicitly at-least-once, not exactly-once, because PROTOCOL.md reserves exactly-once for approvals. It handles out-of-order replay and overlapping batches; a full replay accepts nothing new. WHAT IS STILL NOT DONE, stated plainly: The Durable Object that will host this. Single-threaded execution per tenant IS the isolation boundary, and that is a property of Cloudflare, not of this repo. Without an account there is nothing to deploy to and no way to verify the boundary that matters. Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
ADAPTERS.md opened with "verify, don't assume — this table will rot". This makes that a measurement rather than a recollection: `pnpm probe:runtimes` probes each installed CLI and emits the matrix, with `--json` for the machine-readable form. It is not in CI, because which CLIs are installed is an environment fact — a missing runtime is reported, not failed. The probe settles two things the docs previously only asserted. 1. "Sessions run YOLO by default" is now VERIFIED, not quoted. Four of five CLIs expose a permission-bypass flag and oar passes it: claude --dangerously-skip-permissions, grok --always-approve, kimi --yolo. Only codex documents a sandbox (-s/--sandbox). The premise this entire product is built on is checkable, and it checks out. 2. FOUR OF FIVE RUNTIMES SELF-UPDATE — claude/codex/grok `update`, kimi `upgrade`. `k.managed-copy-never-self-upgrades` (I15) was recorded for k-carrier, but the reasoning is not specific to it: a managed copy that also upgrades itself is a second writer on a component the host believes it owns. If Radius ever manages a copy of any of these four — exactly the k-carrier pattern — their built-in updater is that same second writer. So I15's scope is every managed binary on the host, not one vendored component, and this probe is how the condition is noticed appearing somewhere new. Still unmeasured, and stated as such in the doc and SAFETY.md §3: a documented flag is not a working sandbox. This measures what each CLI CLAIMS to support; the isolation it actually enforces is unknown. Writing the probe caught a stray TypeScript `as` in a .mjs file — caught by pnpm lint, which is the gate doing its job. Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
ROADMAP calls the public API "the real moat", which only holds if versioning is enforced instead of described. PROTOCOL.md §8's rules are implemented here as a checker: parse, classify, and assert the bump. The rule that matters is the one that surprises people: TIGHTENING A DEFAULT IS A MAJOR. Turning a sandbox on by default is breaking for a host built against v1 defaults even though no field changed — a host that silently loses its sandbox must fail loudly, not drift. A tool that diffs field names cannot see that, so classifyRelease works on semantic facts, not shapes. assertVersionBump fails in BOTH directions, which is the part most teams get wrong. Direction one (a breaking change shipped as a minor) breaks your own client, so everyone catches it. Direction two (an additive change shipped as a major) never breaks anyone — it just forces an upgrade for an optional field, until "stable" means nothing. It feels safe precisely because it is quiet. assertCompatible implements "a host may lag the plane by one minor, never a major", including refusing a two-minor lag: "one" means one, and assuming two is how a compatibility promise quietly dies. MUTATION-VERIFIED (5/5) a tightened default downgraded to minor (3 tests fail), a two-minor lag accepted, a major mismatch accepted, an additive change allowed to bump major, and version parsing made permissive. The first attempt at the first mutation left a dangling `||` and produced a syntax error rather than a real failure — a false catch, redone properly. A mutation that fails for the wrong reason has verified nothing. STILL NOT BUILT: the wire transport, an SDK, and the compatibility policy as a published document. This is the version surface those would sit on. 167 tests, 8/8 gates. Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Seven commits that had been sitting on a local branch with no remote. Radius is now published, so this proposes them for review.
radius/v1protocol versioning, enforced rather than promised: a breaking change shipped as a minor is refused, and a version that moves backwards is refusedVerified locally:
pnpm typecheckclean,pnpm testgreen (81/81 in@graycode/protocol),.github/workflows/ci.ymlis already present so this will be gated.