Skip to content

feat: Phase 0 supervisor, e2e harness, and radius/v1 protocol versioning - #1

Merged
Patel230 merged 7 commits into
mainfrom
feat/phase0-supervisor-and-e2e
Sep 29, 2026
Merged

Patel230 merged 7 commits into
mainfrom
feat/phase0-supervisor-and-e2e

Conversation

@Patel230

Copy link
Copy Markdown
Contributor

Seven commits that had been sitting on a local branch with no remote. Radius is now published, so this proposes them for review.

  • radius/v1 protocol versioning, enforced rather than promised: a breaking change shipped as a minor is refused, and a version that moves backwards is refused
  • durability harness and Phase 2 sync logic
  • provider counts derived from the registry where they were hardcoded

Verified locally: pnpm typecheck clean, pnpm test green (81/81 in @graycode/protocol), .github/workflows/ci.yml is already present so this will be gated.

Finishes the two Phase 0 items that were buildable without a product
decision, and records what the resulting evidence does and does not prove.

AgentSupervisor (src/supervisor.ts, 7 tests)
  Three properties, each a way unattended agents go wrong:
  - never two writers on one stream; start() refuses while running
  - nothing observed left unpersisted; stop() drains even when dispose throws
  - no process residency — a wedged harness raises StopTimeoutError rather
    than hanging the host forever

  The residency test spawns a REAL OS process, records its pid, stops, then
  polls until the pid is gone. A mock asserting "dispose was called" would
  pass while the process lived on, which is the bug that property exists for.

kill -9 e2e (test/e2e-kill.test.ts), run against a real Claude session
  10 records durable after a real SIGKILL: contiguous from seq 0, all three
  oar record kinds, replayed identically by a fresh store. The prompt, the
  accepted response, system/init, the model's assistant frame and a terminal
  result/success all survived — stamped with OUR stream name, not oar's native
  id, so the untrusted-name property holds in real conditions.

  Stated plainly: on this machine the turn COMPLETED before the kill, even at
  a 9s delay. So this proves records survive a hard kill; it does NOT yet
  prove a kill *during streaming* is safe. The run now reports which case it
  hit, and RADIUS_E2E_KILL_MS exists to aim at a colder one. Mid-generation
  remains unproven.

Also measured, because D-004 was blocking Phase 0's own gate:
  cargo build --lib succeeds (~11s); cargo build --all-targets FAILS. Five of
  six declared Cargo targets are missing — the whole examples/ directory was
  never vendored. So a path dependency works and a built binary does not.
  check:notices now parses the manifest and fails on an unrecorded missing
  target, so a re-vendor cannot silently drop one.

Gates: 8/8 green. 39 tests.
Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
…nforced

Phase 1.1 and 1.4: the two pieces whose logic can be proved exhaustively
without a platform runtime. packages/protocol, 24 tests.

Principal (1.1)
  The core type, with the deliberate omission made STRUCTURAL rather than
  conventional: no token/secret/apiKey field, and parsePrincipal rejects
  unknown fields so a credential cannot ride in through an untyped JSON bag.
  DATA-MODEL.md §4 asks for "a test that scans for it. Not a code-review
  convention" — so the tests grep serialized principals AND grants for nine
  secret shapes, with a positive control that plants a real-looking AWS key
  and requires detection. A scanner matching nothing would otherwise pass
  every other test and look green.

  CapabilityGrant.expiresAt is required in the type, so a permanent grant is
  unrepresentable rather than merely discouraged (DATA-MODEL §2, D-008).

Lease manager (1.4)
  The primitive oar explicitly declines. Exclusive claims, epoch fencing,
  and expiry computed from process.hrtime — never Date.now(), which the user
  owns and can rewind with one `date` command.

  The restart case is the subtle one and D-008 names it: hrtime resets on
  restart, so restore() adopts a persisted anchor and DROPS any lease whose
  window elapsed while the process was down. Reviving it would let a rewound
  clock resurrect a dead lease.

What this changes about the project's honesty:

  check:invariants now reports "11 specified only, 3 enforced in code,
  1 verified against vendored source" instead of implying all 15 were prose.
  I3 (expired lease refused), I5 (stale epoch rejected) and I14 (clock
  rollback cannot extend a lease) moved from spec to enforced code.

  Each is mutation-verified: removing the epoch check, dropping the expiry
  check, switching expiry to Date.now(), freezing the epoch counter, and
  making the parser permissive each fail the suite.

Two real bugs were found by these tests, not by review: the monotonic
clock's continuity formula was wrong across a restart (time jumped to 0,
which would have expired every live lease on any host bounce), and
LeaseManager had no way to rehydrate leases at all — a Durable Object could
not have restarted.

Gates: 8/8 green. 63 tests.
Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
Phase 1.2, minus the part that cannot honestly be built yet.

WHAT IS BUILT AND TESTED (28 tests)
  TokenStore   tokens scoped to (principal, launchId, capability); 256-bit
               CSPRNG; only sha256(token) is retained, so a leaked table
               yields nothing replayable (DATA-MODEL.md §4). Monotonic
               expiry, launch binding, revocation without agent restart,
               revokeLaunch as a lost-laptop kill switch, and narrowing of a
               live token that can only ever shrink.
  Policy       deny by default. An ungranted capability is refused, and
               every decision is audited INCLUDING denials — a log that
               records only successes shows you nothing when something goes
               wrong. D-014 enforced here as well as in the schema, so a
               caller cannot skip the table. An expired grant reports as
               expired rather than as a generic refusal, because an operator
               debugging needs the real reason.
  Provider     scoped behind a narrow interface. ResolvedCredential is
               opaque — no accessor, no toString — so a raw key is one
               careless JSON.stringify away from a log line otherwise, and
               the only way to use one is to hand it back to the adapter that
               issued it.

WHAT IS NOT BUILT, AND WHY
  Real credential scoping against Anthropic/OpenAI/Google/xAI. That needs
  live provider keys. A broker that looks finished while credentials escape
  is the worst thing this project could ship, so the boundary is isolated and
  every provider path is marked unverified rather than stubbed into looking
  real. Unscopable providers are REFUSED (assertScopable) rather than given a
  raw key — that check is what stops the product quietly regressing to YOLO.

  Still to prove with a live key: that a scoped credential cannot act beyond
  its scope; that revocation stops the upstream call and not just our
  bookkeeping; and that the credential is unreachable from agent context, a
  heap dump, and error messages.

  The loopback HTTP proxy is not built either — it is plumbing around a
  decision engine that is now correct, and building it before the provider
  boundary is settled would put an unverified surface in front of a real key.

MUTATION-VERIFIED (7/7 caught)
  storing the raw token, dropping the monotonic expiry check, making narrow
  widen, allowing unknown capabilities, letting an agent approve itself,
  changing the denial reason, and neutering revokeLaunch.

Gates: 8/8 green. 91 tests.
Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
The remaining locally-buildable work. 127 tests, 8/8 gates.

Sandbox policy (1.3) — deny by default, and honest about measurement
  The design problem I8 poses is "measured, or not claimed", and the way
  most code fails it is by behaving safely while its docs say "sandboxed".
  So `unknown` and `unmeasured` are first-class values, claimsIsolation()
  returns false for both, and an unprobed runtime is REFUSED at decision
  time — not merely described cautiously. A module that reads well and
  reports a sandbox nobody probed is exactly how a claim becomes fiction.

  Unprobed > explicit grant > escalate > deny, and an escalation carries the
  exact request so the required approvals row cannot be lost in translation.

Approvals (D-013) — exactly-once, by construction
  One slot keyed by requestId, so a client retry physically cannot produce a
  second row. A replay carrying DIFFERENT content is refused rather than
  overwritten, because overwriting would let an agent re-approve a decision it
  was never granted. This is the only exactly-once claim in the system, and
  PROTOCOL.md is emphatic it must not be blurred into a general one.

Cost ledger (Phase 4 gate) — a cap that holds under fault injection
  The gate says "cannot be exceeded, including under retry and partial
  failure". Retries cost twice, partial failure splits reservation from
  settlement, and concurrent turns race — so the cap is enforced at RESERVATION
  time against committed+reserved, and reserving the upper bound means a
  partial result cannot overshoot. A repeated settlement is a no-op, a crash
  between reserve and send releases headroom rather than leaking it, and a
  cost above the reservation is refused rather than absorbed.

A real bug the tests caught: reservedFor counted SETTLED reservations, so
headroom was charged twice — available() understated the budget and would have
silently throttled an agent that still had money.

MUTATION-VERIFIED (7/7 caught)
  a retry creating a second row, an agent approving itself, a cap ignoring
  in-flight reservations, a double settlement, an unprobed runtime being
  allowed, claimsIsolation always returning true, and describeIsolation
  claiming a sandbox for an unprobed runtime.

STILL UNVERIFIED, and it stays that way: real provider credential scoping,
and the actual isolation strength of any runtime. The policy is built and
tested; the boundary it guards is not yet demonstrated with a live key or a
real sandbox probe.

Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
Closes the remaining locally-buildable work. 148 tests, 8/8 gates.

Durability harness (milestone 0.4)
  Truncates the durable stream file at EVERY byte offset and asserts what
  survives is always a contiguous prefix [0..k] — no gap, no duplicate, no
  half-written record — then replays a truncated stream and asserts the
  resume delivers every later record exactly once.

  It does not assert "nothing is lost", because that is false and RecordWriter
  documents why: a record observed but not yet fsync'd dies with the process.
  A corrupt prefix is unacceptable; a lost tail is expected. Asserting the
  stronger false thing would have made this harness look better and be worse.

  Writing it caught the harness's own floating promises first — `void
  store.append(...)` raced the teardown. The exact hazard RecordWriter exists
  to prevent, caught by the thing built to detect it.

Phase 2 logic (sync.ts)
  Tenant resolution, outbox-then-settle, and cursor merge — the parts of the
  control plane that are logic rather than deployment.

  Tenant resolution refuses an unknown token rather than defaulting, because a
  default tenant is a silent cross-tenant read. The refusal names why it
  happens at the edge: resolving inside a DO means the DO was already addressed
  by an id the client chose.

  Outbox settle is idempotent in both directions — a double settle is a no-op,
  and re-enqueuing an id does not clobber a row another worker is still
  settling. The property is redelivery, never loss.

  CursorMerger is explicitly at-least-once, not exactly-once, because
  PROTOCOL.md reserves exactly-once for approvals. It handles out-of-order
  replay and overlapping batches; a full replay accepts nothing new.

WHAT IS STILL NOT DONE, stated plainly:
  The Durable Object that will host this. Single-threaded execution per tenant
  IS the isolation boundary, and that is a property of Cloudflare, not of this
  repo. Without an account there is nothing to deploy to and no way to verify
  the boundary that matters.

Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
ADAPTERS.md opened with "verify, don't assume — this table will rot". This
makes that a measurement rather than a recollection: `pnpm probe:runtimes`
probes each installed CLI and emits the matrix, with `--json` for the
machine-readable form. It is not in CI, because which CLIs are installed is
an environment fact — a missing runtime is reported, not failed.

The probe settles two things the docs previously only asserted.

1. "Sessions run YOLO by default" is now VERIFIED, not quoted. Four of five
   CLIs expose a permission-bypass flag and oar passes it:
   claude --dangerously-skip-permissions, grok --always-approve,
   kimi --yolo. Only codex documents a sandbox (-s/--sandbox). The premise
   this entire product is built on is checkable, and it checks out.

2. FOUR OF FIVE RUNTIMES SELF-UPDATE — claude/codex/grok `update`,
   kimi `upgrade`. `k.managed-copy-never-self-upgrades` (I15) was recorded
   for k-carrier, but the reasoning is not specific to it: a managed copy
   that also upgrades itself is a second writer on a component the host
   believes it owns. If Radius ever manages a copy of any of these four —
   exactly the k-carrier pattern — their built-in updater is that same
   second writer. So I15's scope is every managed binary on the host, not
   one vendored component, and this probe is how the condition is noticed
   appearing somewhere new.

Still unmeasured, and stated as such in the doc and SAFETY.md §3: a
documented flag is not a working sandbox. This measures what each CLI
CLAIMS to support; the isolation it actually enforces is unknown.

Writing the probe caught a stray TypeScript `as` in a .mjs file — caught by
pnpm lint, which is the gate doing its job.

Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
ROADMAP calls the public API "the real moat", which only holds if versioning
is enforced instead of described. PROTOCOL.md §8's rules are implemented
here as a checker: parse, classify, and assert the bump.

The rule that matters is the one that surprises people: TIGHTENING A DEFAULT
IS A MAJOR. Turning a sandbox on by default is breaking for a host built
against v1 defaults even though no field changed — a host that silently
loses its sandbox must fail loudly, not drift. A tool that diffs field names
cannot see that, so classifyRelease works on semantic facts, not shapes.

assertVersionBump fails in BOTH directions, which is the part most teams get
wrong. Direction one (a breaking change shipped as a minor) breaks your own
client, so everyone catches it. Direction two (an additive change shipped as
a major) never breaks anyone — it just forces an upgrade for an optional
field, until "stable" means nothing. It feels safe precisely because it is
quiet.

assertCompatible implements "a host may lag the plane by one minor, never a
major", including refusing a two-minor lag: "one" means one, and assuming
two is how a compatibility promise quietly dies.

MUTATION-VERIFIED (5/5)
  a tightened default downgraded to minor (3 tests fail), a two-minor lag
  accepted, a major mismatch accepted, an additive change allowed to bump
  major, and version parsing made permissive. The first attempt at the first
  mutation left a dangling `||` and produced a syntax error rather than a
  real failure — a false catch, redone properly. A mutation that fails for
  the wrong reason has verified nothing.

STILL NOT BUILT: the wire transport, an SDK, and the compatibility policy as
a published document. This is the version surface those would sit on.

167 tests, 8/8 gates.

Signed-off-by: Lakshman Patel <Lakshmanp230@gmail.com>
@Patel230
Patel230 merged commit 19405a9 into main Sep 29, 2026
1 check passed
@Patel230
Patel230 deleted the feat/phase0-supervisor-and-e2e branch September 29, 2026 17:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant