Repository navigation
Conversation
added 30 commits
September 29, 2026 21:41
Stand-in invocations ran concurrently (a foreground `up` while the driver polls `ps`) and each did an unlocked read-modify-write of state.json. A `ps` that loaded the state before the second `up` saved its foreground token then saved its stale copy over it, so `up` saw its token gone and exited 0 before readiness; lost call records failed the acceptance test as well. That failed 26 of 40 local runs, and Runtime state models on CI. Each invocation now holds an exclusive flock across its read-modify-write and releases it only before its long waits (the foreground loop and the ps-hang fault). 40 of 40 local runs pass.
A review of #109 found two startup ownership gaps. The native HTTPS owner discarded the identity listenPublishedUnixSocket returned and only recorded one after lstat, chmod and lstat of the path. A failure after publication then skipped listener and endpoint cleanup, and a replacement in that window could be chmodded or recorded as the endpoint. The owner now keeps the published identity at once, compares its first observation against it, and never changes the endpoint by path. Failure cleanup always closes the listener, which can no longer unlink anything but the retired staging name, and then removes the endpoint only while it is this listener's inode; a replacement is kept. The helper awaited listen() outside its cleanup, so a listen that bound the staging name and then failed could leak that socket and listener. Listening is now inside the cleanup, and a failed start removes only a staging socket this user created. The helper also sets the endpoint mode (0600) on the staging inode before publication, so the MCP backend, the HTTPS owner and the owner challenge no longer chmod the public path. Controls: bind-then-fail, mode-preparation failure, an endpoint at the AF_UNIX path limit (and one byte over: refused cleanly on Bun 1.3.9, bound in full on 1.4.2, never truncated), and an owner fault or replacement between publication and first observation, on Bun 1.3.9 and 1.4.2.
The private-bind fixture hooked chmod on the public mcp.sock to prove the socket is private before its mode is set and that the creation mask is restored. Since the endpoint's mode is now set on its staging socket before publication, the hook never fired and the check failed. It now observes that staging chmod with the same privacy and mask checks, requires mode 0600, and fails if the published endpoint is ever chmodded by path.
…a socket A second review of the Unix-socket publication found three staging-path windows where an entry this attempt could not prove was its own could be changed or removed: - After an ambiguous partial bind (no recorded identity), cleanup unlinked any same-uid socket at the staging name. - The mode was set by path after an awaited lstat, so a replacement in that gap could be chmodded before the postcheck refused it. - On success the staging name was unlinked unconditionally after the awaited link and endpoint lstat. The socket is now created with exactly its mode (the umask during the synchronous bind), so nothing is ever chmodded, and its identity is recorded in the same tick. Publication links only while the staging name is still that socket, and retirement removes it only while it is (check and removal back to back). An entry that is not this socket, or cannot be proven to be, is never removed: on failure it is moved aside while the server closes, so the runtime's close-time unlink by name cannot reach it, and then put back with the same inode. The owner challenge's chmod dependency becomes an afterOwnerSocketPublish test seam, and the private-bind fixture now requires the socket to be created private with mode 0600 and no socket name to be chmodded.
…ng entry The failure close moved an unproven staging entry to one random holding name: if that name was occupied the close went ahead unprotected, and a holding entry created between the check and the rename was overwritten. The entry is now moved with link(2), which never replaces a holding entry, over a bounded list of fresh holding names, and the staging name is dropped only while it is still the linked entry. An entry that cannot be moved aside (every holding name occupied, or not hard-linkable) leaves the close to proceed, now documented as a residual. The helper's documentation now states its scope: accidental and concurrent entries in a private directory, not an adversarial process of the same user, with the remaining path-based and post-return residuals.
added 5 commits
September 30, 2026 15:26
Add explicit selector-bound recovery for a dead legacy HTTPS owner beside an unpublished shared-owner configuration. Preserve original inodes and the CA with resumable exclusive archival, held IPv4/IPv6 guards, and refusal on live or uncertain evidence.
Keep strict receipt identities as the default. Allow only an explicitly selected source-device witness and both current HTTPS inodes, with unchanged inode numbers and matching old/current device mapping. Recheck current owner, host and guest boot, source share and stopped graph scope before archival; preserve historical bytes and report original host-volume continuity as unproven.
This was referenced Sep 30, 2026
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Explicit frontend recovery could fail before cleanup when native recovery had already stopped a graph but the frontend had not saved a restart intent. Preflight also resolved mutable image tags again, potentially reviewing a different image than retained startup would actually use.
Recovery preflight now selects stopped review only after fresh native inspection confirms the exact run/owner/namespace/plan, complete journal, absent recorded compute and present recorded volumes, then repeats that proof after review. For unchanged original Compose input it uses the same authenticated retained image selector as startup. Changed original input keeps normal image resolution and compatibility checks. Active graphs retain service authority and live dependency-listener requirements. Normal intent, retaining down, frontend recovery, publisher retirement, finalization wait and startup sequencing remain intact.
Validation: 26 focused restart tests, including stopped recovery without an intent, incomplete/foreign observations, drift after review, active listener requirements, retained image reuse versus changed-input resolution, and malformed provenance refusal. The final full CLI gates passed: typecheck and check, plus 1,737 Bun tests passed, 67 skipped and zero failures. Rust is unchanged from #119's passing default/all-feature suites and strict Clippy.
A live read-only M3 preflight confirms that using admitted image IDs restores the exact recorded normalized input: zero image, route or port changes, with source compatibility passing. Normal signed CLI restart passes this preflight and reaches publisher recovery. It currently refuses when the generic publisher retirement command selects an old crash-recovery receipt after a later ordinary acknowledged shutdown; the separate bounded correction is PR #121. Application/browser acceptance remains open.
Exact-head CI: seven jobs passed; the macOS general test job crashed in Bun 1.3.9 (segmentation fault, exit 133). The one failed-job rerun passed; all eight exact-head checks now pass. The runtime crash is tracked separately in HACK-1211 and the Bun 1.4.2 qualification.
Targets
next; includes current recovery integration and admission fix from #118/#119. No application repository edits and no publication. This component evidence does not establish whole-product parity or prerelease readiness.