docs(adr): land ADR 0158 — the silent-controls taxonomy the ledger already pointed at - #145
Merged
Conversation
Every coordination control in this repo is PULL-based: a new session discovers its peers from the SessionStart banner and the peers learn nothing until someone trips the collision gate. That is too late for the collision that costs the most -- two sessions building the same THING in different files, where nothing file-shaped can catch it. This closes the push direction. It ASKS, it cannot send. Hooks are shell commands and session messaging is MCP, so the hook prints the instruction, the live peer roster and the id-resolution rule at the first prompt that has intent to report; the model does the sending. UserPromptSubmit, not SessionStart: at SessionStart a session knows it exists and nothing else, so it can only say hello -- the interrupt without the information. THE ID RULE IS THE PAYLOAD, and it is counter-intuitive enough that the text states it with its evidence. The registry id in this repo's banners is NOT the MCP session id; measured, a registry id and an MCP id for one session shared no characters. Branch does not join them either -- the two rosters reported different branches for the same checkout in 2 of 6 cases. Only cwd joins, and it must be matched EXACTLY: every worktree cwd is an extension of the primary's, so a prefix match resolves a peer in the primary to an arbitrary worktree session. A registry id passed to send_message fails SILENTLY, which reads as the peer ignoring you. EVERY DECISION LEAVES A RECEIPT, because the bug being fixed was a hook that was wired, fired, resolved nothing and exited 0 for weeks -- byte-identical to a healthy hook with no peers. For the same reason the shim carries its OWN missing-script notice: every receipt the hook writes lives INSIDE the script, strictly downstream of the resolution failure that IS the bug, so the shim is the one surface that still reports when the script does not resolve. It is gated on presence.ps1 so the entry stays silent in every unrelated repo on the machine. It always exits 0 -- a UserPromptSubmit hook that fails can block the user's prompt. It consumes presence.ps1 and therefore the single liveness fence; it does not invent a second notion of live. A separate 'mefor-announce' marker keeps it outside install-coordination's mefor-coord strip and outside the website repo's mefor-web-announce entry in the same settings file, so no installer can delete another's hook, and -Only UserPromptSubmit -Uninstall removes announce alone without disarming the collision gate.
Most tests for a hook like this assert an ABSENCE, and a hook that does nothing at all satisfies every one of them -- which is precisely the production failure being fixed. So the silence assertions are paired with a positive arm: two tests run the SAME runner against fixtures differing only in whether a peer exists, and if the silence tests ever start passing for the wrong reason the positive one goes red first. test_announce_wiring.py is the class the repo had no test for AT ALL: does the thing that gets INSTALLED reach a script that EXISTS, and does it say so when it does not? Its absence is exactly how a wired-but-inert shim survived for weeks. test_every_wired_script_exists_in_this_checkout was written FIRST and watched fail, naming the missing script and printing all three paths it scanned; a green gate is only evidence if it was shown it can see the failure. Also pinned, each because it was got wrong somewhere first: - The foreign UserPromptSubmit entries -- another repo's shim and an unmarked waiting-flag cleanup -- survive install AND uninstall byte-identical. That is the only thing standing between a one-line wiring edit and deleting a hook this repo does not own. - A peer with no StartedAt ranks LAST, not first. ConvertFrom-Json coerces ISO-8601 to DateTime while the '' fallback stays String; Sort-Object over that mixed column raises ZERO errors and puts the empty string FIRST, so without an explicit projected key the least-trustworthy row silently takes the top of a capped target list. - NO_SESSION_ID and DISABLED write their receipt with NO injected -StateDir. An earlier draft resolved the state dir after those branches, so the receipt was unwritable in production while a test that always injected one went green. - Self is excluded by BOTH nets independently: a roster that cannot tell you from a sibling makes the session message itself. - Hostile peer text cannot escape the peer-data block or emit a non-ASCII byte, a hostile session id cannot escape the state dir, and two ids that sanitise identically get two markers. - Two concurrent runs announce exactly once. session-context.ps1 is registered twice on this box today, so double firing is a live pattern, not a hypothetical.
…about .claude WORKTREES.md gains the "Announcing yourself" section that the hook's own emitted text and the shim's missing-script notice both cite by name, so the pointer has to land on main in the same merge. It states the id rule ONCE, as the source of record: registry id is not the MCP id, cwd is the only join key and must be matched exactly rather than by prefix, a usable id starts with local_, and a wrong one fails silently. It also states what the change does NOT do. There is no receive-side hook, so the rule that an announcement is peer DATA -- not an operator instruction, and not something to reply to -- lives in the prose and in the fixed message shape and nowhere else. Reachability is given honestly: presence.ps1 is authoritative for who EXISTS, list_sessions only for who can be MESSAGED, and measured, they disagreed 6-to-1. Cost is stated rather than left to be discovered. CORRECTION, and it is why this doc change is in scope rather than deferred: the same chapter claimed ".claude/settings.json is tracked (shared across worktrees)". It is not. /.claude/ is git-ignored, and git ls-files .claude/ returns nothing -- so a worktree's copy is a creation-time snapshot nothing refreshes and several siblings have none at all. That sentence sat at the exact point a reader decides where to install a hook, and it argues for the wrong answer; the new section directly contradicted it. SESSION-DRIFT-CONTROLS.md records announce as the only PUSH control in the D4 layer, plus the two new guarantees worth tracking separately: that wiring reaches a script that exists, and that a resolution failure is now reported by the shim.
…nd finished
Reported by another session with a repro: it committed a file, went clean, said
in writing it was done and handed the file over -- and the peer it handed off to
was still refused the edit.
overlap.ps1's `Files` is the UNION of what a branch COMMITTED-and-not-yet-landed
with what is dirty in its tree. The gate denied on any live row in that set, so
"this branch authored it" was treated as "someone is typing in it right now".
Those are different claims. The first stays true for the branch's whole life;
only the second is what the gate exists to detect.
It self-clears on merge -- overlap already intersects three-dot with two-dot so a
LANDED branch stops claiming its files. But nothing clears it before landing, and
with PRs currently unable to merge, "until it lands" is indefinite: the blocked
set grows monotonically and is never released. Two sessions that coordinated
correctly and explicitly still cannot hand a file over. That is precisely the
failure this gate's own docstring names -- a gate that cries wolf gets
uninstalled.
overlap.ps1 already told callers to treat its signals differently ("block on
live, mention dormant"), but no caller COULD: the row unioned the two signals
away. So the row now carries `Dirty`, and the single-file query sets
`MatchedDirty` saying which signal actually matched.
The gate now DENIES only on an uncommitted edit in a live worktree, and REPORTS
committed-and-clean as context instead -- the peer may already have done what you
are about to do, which is worth knowing and not worth refusing over.
Fails SAFE across the upgrade: a cached row predating `MatchedDirty` has no such
property and is treated as dirty, so the gate degrades to its previous
over-blocking rather than silently permitting a real collision.
Also, while in the file: `git status` now runs with --no-optional-locks. A plain
status REWRITES the index of the repo it inspects, and this walks every peer
worktree -- so merely asking "what is in flight" was mutating other sessions'
checkouts.
Verified against the live repro and both directions: the reported file now
allows with context; a file with uncommitted changes in a live worktree still
denies; an untouched file stays silent.
…rktree name Reported by the session it happened to: its worktree is named inter-session-communication-*, auto-generated at creation from a task that session has never worked on -- it has been doing ASVS scorecard work for its entire life. The directory name is the most visible identifier in presence.ps1, overlap.ps1 and this hook's output, and it had already misled TWO sessions (including this one) into guessing that session was building the announce hook. A worktree name is a creation-time label, not a statement of current work, and nothing keeps the two in sync. The claim note is the only field written DELIBERATELY to say what a session is doing, so the roster now prints it, and the legend tells the reader to prefer it over the name. Joined on the claim's `worktree` path, normalised the same way as every other cwd key here. Fail-open throughout: no claims directory, an unreadable claim, or a peer with no claim all just mean the name is the only thing we have -- which is exactly the status quo, never an error. Same session also flagged that the branch I read for it from list_sessions was stale (a spent, merged branch). The announce text already refuses to join on branch and says why; this is a second, independent reason not to trust it.
A control that cannot distinguish 'ran and resolved' from 'ran and found nothing' is not installed, however it looks. The announce shim outlived every other silent-control defect found the same day BECAUSE it printed a status message -- which is more convincing than silence. The structural cause is the reusable part: every receipt that hook would have written lived inside the script the shim failed to find, so every check sat strictly downstream of the failure it existed to detect. Looking was not neglected, it was impossible. The question to ask of a new control is which surface still reports when the control itself fails to load. Formulation owed to a peer session that hit four instances of this class in one day and named it more sharply than I had.
…e hard way Announce-on-join introduces a session; it does not let an established one push an operational notice. That increment is deferred, and on 2026-08-01 six sessions rehearsed it by hand for four hours. Three constraints fell out, recorded so the next attempt does not rediscover them: - A broadcast needs an EXPIRY or a predicate the RECIPIENT can evaluate, never a promise from the sender. A merge freeze shipped with 'lift when #119 merges'; #119 died on an unrelated CI timeout, so five sessions held on a condition that could not arrive and a second round was needed to retract it. - 'Don't do X' is the wrong primitive when automation already has X armed. The freeze asked for restraint while six PRs had auto-merge ARMED and would have landed with nobody clicking anything. The right ask was an action: disarm. - Coordination a tool cannot read does not count. Two sessions agreed IN WRITING to hand over a file and the gate still refused, because the agreement was prose and the gate reads git. Field data from the sessions that lived it, not speculation.
Nothing drove overlap.ps1's row computation against a real repository, so the question "does MatchedDirty hold when a file is dirty AND committed at once" was unanswerable by the suite. Raised by the session that spent an evening in exactly that state. THAT CASE IS THE ONE THAT FAILS SILENT, which is why it gets a real fixture rather than a stub row. A peer with uncommitted edits in one region and landed work in another is a genuine collision. Had MatchedDirty been derived from the committed diff instead of the working tree it would read FALSE there, the gate would allow, and two sessions would write one file with nothing reported. The over-block this replaced was loud and annoying; that would be quiet and cost someone their work. Verified the tests can SEE it rather than assuming: sabotaged the row to publish an empty Dirty set -- the precise mis-implementation warned about -- and both MatchedDirty assertions went red; restored, all five green. A test written after the code, never observed failing, is a test of nothing. Also pins that overlap does not rewrite a peer worktree's git index, by comparing the index mtime across two queries. An observer must not perturb what it observes, and this one was doing so on every PreToolUse before f55d6c6. Stub rows would only have asserted that the plumbing carries a value someone else computed; the whole question here is what git actually reports.
…at exists Raised by the session that traced the shim: the coordination hooks are not installed copies, they are inline commands that locate their script in a working tree at every invocation. If neither base yields the file, Test-Path fails, the loop ends, nothing runs, and the tool call proceeds with no hook and no signal. "The hook is uninstalled" and "the hook ran and permitted this" are indistinguishable from outside, and nothing was watching. Not hypothetical: a foreign UserPromptSubmit entry sat in this same settings file for weeks probing a script that exists only in another repo. The risk composes badly for collision_gate.ps1 specifically, which now (a) fails OPEN on any error, (b) denies less by design after the dirty-vs-committed split, and (c) silently no-ops when unresolvable. Individually defensible; together the realistic bad day is "the gate was never running and nobody noticed". This closes (c) -- the observation is not mine, and it is a good one. Found immediately on writing it: FIVE user settings files across account directories, not the one I knew about. The informational test also prints the original defect as output rather than leaving it invisible: FOREIGN UserPromptSubmit [mefor-web-announce] -> scripts/hooks/announce.ps1: RESOLVES NOTHING HERE It is another repo's entry, so this reports it and does not touch it. Carries a NEGATIVE CONTROL, because the assertion passed on the first run and a green that has never been shown to fail is not evidence. The real hooks cannot be unwired to prove the predicate works -- the primary checkout is shared with live sessions -- so it is exercised against a path known not to exist. Local-machine only: CI has no user settings and these skip there, which means CI does NOT guard this property. Said plainly, and every test prints what it scanned BEFORE it can skip, per test_gate_installed_parity.py -- the pytest config has no -rs, so a skip would otherwise render as a bare dot with no reason.
…ommunication-hooks-a52335
…ommunication-hooks-a52335
…ommunication-hooks-a52335
…ommunication-hooks-a52335
Session ended on an owner stop-work instruction at 96% weekly account usage, so this lands the two things that would otherwise have existed only in a transcript. ADR 0158 records a defect class that recurred at least a dozen times across independent surfaces in one working day, in at least two sub-classes: a bound stated independently of the thing it bounds, and a control that cannot observe or act on its own failure. Its spine is that a signal carrying too little information to act on makes every reader re-derive significance by hand until one of them derives it wrong -- so a correct-but-useless RED costs what a silent green costs. EVERY FIGURE IN IT WAS RE-DERIVED BY SOMEONE WHO DID NOT PRODUCE IT, against the repository and the GitHub API. That pass refuted six claims, including four CI numbers that were already merged, and including corrections this session had itself issued hours earlier. Seven retractions are recorded INSIDE the document, each carrying a found-by tag -- because the central empirical finding is that no retraction was made by the author of the claim it retracts, and that is invisible if attribution is smoothed into one voice. Shape over detection is reported as a ratio rather than flattered: three fixes are covered by tests in required CI legs, two by tests that always skip in CI, one by a workflow change with a live residual, and the rest are corrected prose or still open. The Decision separates ENFORCED rules, each naming its gate, from CONVENTION that is knowingly re-breakable. The handoff records what is pushed, what is filed-not-built, and the traps -- a linked worktree's .git being a FILE, a Windows Python unable to read MSYS paths, a raw hasher giving a false mismatch against a git blob on CRLF, and claim.ps1 silently discarding a note refresh. Each is stated as a fact plus its measurement. It also records, first, the five claims this session got wrong -- including retracting a CORRECT estimate on the strength of an incorrect measurement, and sending that false claim to four sessions and the correction to only three. One more arrived while committing this: the leak gate rejected the handoff for a branch slug, on a line a standalone run of the same scanner had passed. The hook scans STAGED files; the standalone run scanned tracked ones. Two scopes, one tool, and only the fail-closed gate could see it. Recorded in the handoff. No engine behaviour changes.
…nothing ADR 0158 was authored and committed in 994bfb1 on claude/intersession-communication-hooks-a52335, a trailing commit pushed about an hour and a half AFTER that branch's PR (#133) had already squash-merged. It therefore never reached main and no PR carried it, while the coordination ledger had already allocated the number: docs/adr/README.md stopped at 0156 and 0158 was taken, so the index pointed at a document that did not exist. That gap had a cost. At least four sessions cited this silent-controls taxonomy as "ADR 0157" -- an unrelated HA demotion-safety document allocated to another worktree and still in flight on PR #139. The document that settles the citation was the one sitting unmerged. This branch is cut from 994bfb1 itself, so the original commit stays in history and authorship is exact. The prose, voice and ASCII-only convention are its author's. This commit drops the session handoff and makes three factual corrections where main moved underneath the branch after it was written, each tagged inline in the ADR's own update convention rather than silently rewritten: * 0fdc326 is unreachable from main (this repo squash-merges). It is now given as "merged as 851c849 (#130)", matching the mapping the ADR already uses for 7ebb2ff/2a6649fb. * transports/email.py and transports/direct.py were cited as carrying the same bare starttls() call. 093db33 (#132) gave both an explicit verifying context; pipeline/alert_sinks.py:384 is now the only remaining instance. * The "the false sentence is still there" claim (five sites, one of them numbered Decision rule 13) is closed out: on main the clause survives only inside its own CORRECTED block at :5270 and as a quotation at :7476. The interval is recorded; the rule it produced is unchanged. HANDOFF-announce-hook.md from 994bfb1 is deliberately not landed: it is session state rather than project documentation, no root HANDOFF-*.md has ever existed on main, and it would publish local shim mechanics into a public repo. It stays on its own branch. Verified: exactly one commit in the repository ever added a 0158 ADR and exactly one 0158 filename exists across all refs, so nothing competes for the number. The index row is unchanged from 994bfb1 and appears exactly once. No engine behaviour changes.
This branch is cut from 994bfb1 so the original commit stays in history and authorship is exact. That commit predates a dozen merged PRs, and its own branch reached main as the #133 SQUASH (3389aa2), so the pre-squash ancestors are not main's ancestors. Merging main back in is what keeps this a two-file change instead of a revert of everything that landed since. Five files conflicted and main's side was taken for all of them -- docs/SESSION-DRIFT-CONTROLS.md, docs/WORKTREES.md, scripts/hooks/announce-session.ps1, tests/test_collision_gate.py, tests/test_coord_overlap_signals.py. Every one is #133 work that is already on main via the squash, plus later fixes; this branch has nothing to add to any of them, and the still-open #140 owns the next round. Verified after resolution: `git diff origin/main` is exactly two files -- docs/adr/0158-*.md and its one index row -- 475 insertions, 0 deletions.
The correction I added said 093db33 (#132) left alert_sinks.py as "the only remaining instance on main". That is a checklist-shaped claim with an expiry date: BACKLOG #323 layer 3 (PR #142) closes the alerts call site, and the sentence goes false the moment it lands. Dating the observation does not help a reader who greps for it in a month and finds nothing. Restated as what happened rather than what is currently true -- #132 closed the two connectors, the alerts call site is tracked as #323 layer 3 -- so it holds whether or not #142 merges, and it says outright that the current state must be grepped rather than cited from here. Deliberately does NOT assert that #142 closed the cell: #142 is open at time of writing, and asserting a merge that has not happened is the same defect pointing the other way. found by: the repo-security-review session, which owns #142 and re-derived all three call sites against origin/main before raising it.
wshallwshall
added a commit
that referenced
this pull request
Aug 2, 2026
…at hid in the analysis Three sessions reviewed this; each correction below is theirs, verified here rather than adopted. RATES. Recomputed exactly: 1 in 1,821 per assertion per leg, 1 in 911 per leg, 1 in 304 per full CI run (I had floored two of them). JANE at test_content_search.py:123 is 8.58e-6 = 1 in 116,509. THE ZERO. The 14-char row was written as "p=0 (unreachable)". It is 7.44e-24. Cause reproduced: 1-(1-64**-14)**144 UNDERFLOWS to exactly 0.0 in float64, silently, with no warning, in a column of plausible values -- inside an analysis arguing that token length is the discriminator, at the one row where length breaks the arithmetic. The item now carries the trap, the exact value, and the idiom that does not underflow (-expm1(N*log1p(-x))). Every figure recomputed both ways and agreeing. THE COUNT. Three different numbers were quoted before anyone checked (7, then 5, then ~16 for a different denominator). Re-derived from the tree: 7 lines / 8 clauses against ciphertext in test_store_encryption.py, 2 of them unsafe; ~16 repo-wide. The item now states the basis AND the exclusions (:905-908 and :927-928 are caplog assertions against log text, not ciphertext; :56/:532/:625 are full-plaintext; :512 is deterministic twice over) so the count is not re-litigated a fourth time. #344 CITATION. Kept -- its owner confirmed the framing and supplied the discriminator that stops a reader folding the two: this one would fire at exactly the same rate on an infinitely fast machine. Related line now says what the citation is NOT. #346 / ADR 0158. Added as the closer sibling: an assertion that passes for a reason unrelated to the property it tests. ADR 0158 cited WITHOUT a link -- I guessed its filename, checked, and was wrong; the file is on PR #145's branch, not on main. VERIFICATION. backlog_status_check.py falsified against this item: a deliberately doubled banner makes it fail at BACKLOG.md:8429 naming #347. The first probe attempt silently no-op'd on a cp1252 decode and "passed" -- the same shape this item is about.
wshallwshall
added a commit
that referenced
this pull request
Aug 2, 2026
…mber The session landing ADR 0158 (PR #145) confirmed #347 is a genuine instance and supplied the precise anchor instead of a general pointer. Verified against the ADR on its branch before citing -- all three lines read verbatim as quoted: :246 "An equality check satisfiable by coincidence is not an equality check." :60 Class 2 -- a control that cannot OBSERVE or ACT ON its own failure :61 Test: "if this control were broken, what would tell me?" :56 / :439 the taxonomy explicitly disclaims completeness Citing the RULE is what makes the reference survive renumbering, and the Class 2 test is the sharper statement of this defect than anything I had written: if the encryption were replaced tomorrow with a weak encoding, "DOE" not in raw would still go green. The answer to "what would tell me" is the control itself. Still deliberately unlinked -- 0158 is on #145's branch, absent from main, so a relative link renders broken. The follow-up (file this against 0158 once it is on main) is recorded as NOT done here, with the reason: that session declined to add instances its author did not choose, and padding a rescued document at merge time is its own defect. Their call, recorded so it does not read as an oversight.
wshallwshall
added a commit
that referenced
this pull request
Aug 2, 2026
…ck does the work Two corrections, both prompted by peer review of the section this PR adds. 1. The armed-auto-merge bullet was measured when `allow_update_branch` was `false` on this repo. It was set `true` later the same day, which may have falsified it. Deliberately NOT rewritten: GitHub's documentation does not connect that setting to base-move auto-update, and no back-fill has been observed by anyone -- replacing a stale-but-measured claim with a plausible-but-never-observed one is a strict downgrade. The bullet now carries its measurement date, states the new behaviour is unverified, and asserts nothing about back-fill. Rewrite it when someone records one. 2. The new subsection presented `--is-ancestor` and the two-dot diff as two co-equal checks. They are not. Once `--is-ancestor` passes, the merge base IS origin/main, so two-dot and three-dot compute the same thing and cannot disagree -- the diff is confirmation, not detection, and the trap only exists in the window where that check fails. The table is measured on this branch minutes apart, when a stale local checkout put it on the wrong side of the very trap it documents: three-dot reported 1 file / 50 insertions while two-dot reported 2 files / 52 insertions and 22 DELETIONS. After syncing, both read 1 file / 50 insertions. That is the failure mode a reader is most likely to miss, because on the branch they are most likely to test -- their own, up to date -- the diff agrees with itself. found by: the coordinator session for the dating (its own config change caused the drift, and it escalated rather than rewrote), and the cranky-lumiere session for the load-bearing-check distinction, while independently verifying #145.
…strings The document's own rule, applied to itself: a quoted string survives a file edit, a line number does not. Ten citations replaced. WHY NOW. All three ci.yml citations (:229, :233, :254) resolve to unrelated text the moment #138 lands, and six docs/BACKLOG.md citations had ALREADY rotted on main before that -- +14 to +40 lines of drift from #345/#346/#347 being appended, with every cited claim surviving verbatim at a new address. Measured fresh against origin/main and against #138's branch, not reused from the report that found them. TENSE, not just addresses. Two of the quoted strings do not survive #138 -- "Measured over the 11 PASSING windows-2025 runs" and "1.46x" are both deleted by it, because #138 ADOPTS this ADR's retractions 1-3 wholesale (12:31, 21:34, 25:51, 1.006x, 1.206x, pools 42/39/36). Left in the present tense those two sentences would ship knowingly false the hour #138 merges, so they now say what ci.yml stated when this was written. The retractions themselves are unchanged and are vindicated by #138, not contradicted. ANCHORS ARE SINGLE-LINE ON PURPOSE. A first pass rewrapped two quotes across a newline, which makes them ungreppable and would have swapped one rot for another. Every anchor is now verified to grep as one line AND to resolve in the tree it points at -- "ZERO tests failing" resolves in ci.yml both on main and after #138. pyproject.toml:266 was simply wrong: the zizmor pin is at :271, in the group opening at :268. Replaced with the group name, which is what the sentence needed and cannot rot. The residual it reports -- that the pin's home is outside zizmor's paths filter -- is verified TRUE and unchanged. OUT OF SCOPE, deliberately: line numbers into less volatile files remain (test_stage_dispatcher.py, claim.ps1, zizmor.yml, install-coordination.ps1, freethread-smoke.yml, collision_gate.ps1). So the ADR does not yet "state no line numbers" outright -- see the handoff note.
wshallwshall
enabled auto-merge (squash)
August 2, 2026 18:14
wshallwshall
added a commit
that referenced
this pull request
Aug 2, 2026
…ason (#147) * backlog(#347): PHI-at-rest tests assert short-substring absence against base64 ciphertext `tests/test_store_encryption.py:95` asserts `"DOE" not in raw` against a value encrypted under a fresh random key, so the base64 body is fresh random text every run. Measured p ~ 5e-4 per run per leg; it fired on PR #142's windows-2022 py3.14 leg with encryption working correctly. Filed rather than fixed: the fix direction is the maintainer's call (decoded-bytes assertion vs. non-recoverability vs. full-plaintext absence), and simply widening or deleting the substring check would drop the PHI-at-rest property it reaches for. Includes the sibling audit: the identical 3-char shape survives at :303, a 4-char instance at test_content_search.py:123, and ~13 >=6-char instances whose rate is immaterial but whose shape is the same. `test_off_by_default_stores_plaintext` does NOT share the shape (deterministic equality) and needs no change. * backlog(#347): lead with the instrument, and correct the rate to the per-CI-run figure Reframed after the session that owns PR #142 reproduced the diagnosis independently. Three substantive corrections, none of them cosmetic: 1. FRAMING. The banner led with the flake; it now leads with what is actually wrong. The substring check fails when encryption worked AND would pass on a weak encoding that happened to avoid those three characters. The second half is the PHI defect; the flake is only what made someone look. 2. RATE. 5.5e-4 is per assertion per leg. There are two such assertions and three OS legs (ubuntu + windows-2022 + windows-2025, verified against ci.yml), so the figure an operator experiences is 1 in 303 full CI runs, not 1 in 1,820. Both over-estimate caveats kept: L is from one measured value, and the hex-fingerprint prefix is not base64. 3. SCOPE. Reversed my own recommendation to sweep the >=6-char sites for shape. Rewriting a dozen correct assertions costs review attention for no risk reduction; the pattern-propagation concern is answered by writing the ">=6 characters" rule into the :49-58 convention comment instead. Fix list is now :95, :303, and test_content_search.py:123 (4 chars, unsafe under that same rule -- an addition to the reviewing session's list, in its own framing). Also records the confirmation-by-prediction: #142's re-run came back 25 passed with the prediction written beforehand, so a green re-run confirms a chance collision rather than resetting the question. * backlog(#347): exact rates, a verified count, and the float64 zero that hid in the analysis Three sessions reviewed this; each correction below is theirs, verified here rather than adopted. RATES. Recomputed exactly: 1 in 1,821 per assertion per leg, 1 in 911 per leg, 1 in 304 per full CI run (I had floored two of them). JANE at test_content_search.py:123 is 8.58e-6 = 1 in 116,509. THE ZERO. The 14-char row was written as "p=0 (unreachable)". It is 7.44e-24. Cause reproduced: 1-(1-64**-14)**144 UNDERFLOWS to exactly 0.0 in float64, silently, with no warning, in a column of plausible values -- inside an analysis arguing that token length is the discriminator, at the one row where length breaks the arithmetic. The item now carries the trap, the exact value, and the idiom that does not underflow (-expm1(N*log1p(-x))). Every figure recomputed both ways and agreeing. THE COUNT. Three different numbers were quoted before anyone checked (7, then 5, then ~16 for a different denominator). Re-derived from the tree: 7 lines / 8 clauses against ciphertext in test_store_encryption.py, 2 of them unsafe; ~16 repo-wide. The item now states the basis AND the exclusions (:905-908 and :927-928 are caplog assertions against log text, not ciphertext; :56/:532/:625 are full-plaintext; :512 is deterministic twice over) so the count is not re-litigated a fourth time. #344 CITATION. Kept -- its owner confirmed the framing and supplied the discriminator that stops a reader folding the two: this one would fire at exactly the same rate on an infinitely fast machine. Related line now says what the citation is NOT. #346 / ADR 0158. Added as the closer sibling: an assertion that passes for a reason unrelated to the property it tests. ADR 0158 cited WITHOUT a link -- I guessed its filename, checked, and was wrong; the file is on PR #145's branch, not on main. VERIFICATION. backlog_status_check.py falsified against this item: a deliberately doubled banner makes it fail at BACKLOG.md:8429 naming #347. The first probe attempt silently no-op'd on a cp1252 decode and "passed" -- the same shape this item is about. * backlog(#347): two safety grounds, and a correction to the one I had just adopted The #142-owning session pointed out that :512 is safe for a reason unlike the others -- ':' is not in the base64 alphabet -- and suggested it as a third safety category. It is a real category and it is now ground (1), stated more usefully than either of us first had it: a token containing a character the value cannot contain is a PROOF at any length, and it is the same principle that makes the recommended fix (assert ADT not in raw) deterministic. Rule and remedy are now one idea. But the argument as given does not survive, and I checked before adopting it. The haystack is <marker>:<base64>, NOT base64 alone, and the marker carries colons -- so ':' IS representable and ground (1) does not cover :512. It is deterministic for a different reason: fixed marker layout with the version field reading v1, plus a body with no ':' for the run to straddle. Structure, not alphabet. The item now says so explicitly, because a rule that is right about the conclusion and wrong about the mechanism is the thing this whole item is about. Verified mechanically rather than argued: over the haystack's actual character set, the ADT fixture carries \r & . \ ^ | (unrepresentable -> ground 1 genuinely holds), while DOE, JANE, SECRET, WESTWING, SECRETSTATEMRN and ':v2:' are ALL fully representable -- so ground (1) applies to none of them and length is their only defense. That is what the table already claimed; now it is checked. Also records why the float64 trap survives review, from #344's owner: the naive expression is correct everywhere you would sanity-check it and silently wrong only in the tail. * backlog(#347): cite ADR 0158 by its rule and its class, not by its number The session landing ADR 0158 (PR #145) confirmed #347 is a genuine instance and supplied the precise anchor instead of a general pointer. Verified against the ADR on its branch before citing -- all three lines read verbatim as quoted: :246 "An equality check satisfiable by coincidence is not an equality check." :60 Class 2 -- a control that cannot OBSERVE or ACT ON its own failure :61 Test: "if this control were broken, what would tell me?" :56 / :439 the taxonomy explicitly disclaims completeness Citing the RULE is what makes the reference survive renumbering, and the Class 2 test is the sharper statement of this defect than anything I had written: if the encryption were replaced tomorrow with a weak encoding, "DOE" not in raw would still go green. The answer to "what would tell me" is the control itself. Still deliberately unlinked -- 0158 is on #145's branch, absent from main, so a relative link renders broken. The follow-up (file this against 0158 once it is on main) is recorded as NOT done here, with the reason: that session declined to add instances its author did not choose, and padding a rescued document at merge time is its own defect. Their call, recorded so it does not read as an oversight. * backlog(#347): itemise the provenance, and source the one number that was not mine Three fixes, all of the same defect the item is about -- an observational claim carrying more confidence than its sourcing supports. 1. THE 200k SIMULATION IS NOT MINE. It arrived with the originating defect report and I never ran it; the sentence read as though this filing corroborated the rate that way. Every word was accurate, which is the shape: an unsourced observational claim inside a sentence whose whole job is telling the reader how much to trust the number beside it. Now attributed, with this filing's actual contribution (exact Fraction derivation, cross-checked against expm1/log1p) stated separately. 2. "PRODUCED INDEPENDENTLY BY THREE SESSIONS" was an aggregate confidence claim. Two sessions derived rates; the third contributed process discipline. Replaced with itemised attribution -- who supplied the framing, the >=6 rule, the discriminator, the demand to falsify the gate -- and an explicit statement that no claim rests on a count of who agreed. A session count is not evidence. 3. "#344 IS a fixed bound meeting variable latency" -> "#344's THESIS is". That item's instance 2 has since been re-diagnosed as a swallowed lock-timeout (SET LOCK_TIMEOUT 0 -> native 1222 caught and returned as a normal empty, with the dispatcher then parking in a terminal IDLE) -- not a bound at all. The wholesale characterisation was over-broad, and the item now says not to lean on "#344 = timeouts" as a premise. The chain that prompted this went two sessions -> one -> none -> mechanism-only -> mechanism-only-labelled-as-deduction, on a separate claim, every step a good-faith correction, and the conclusion correct throughout. Only the stated mechanism was hollow, and the stated mechanism is what the next reader carries. * backlog(#347): require the replacement assertion to be falsified before it is trusted The item told an implementer HOW to fix the assertion but not how to know the fix works. Shipping the replacement on an unfalsified green would reproduce the defect inside the remedy -- a green taken as evidence for a property it cannot see is the whole item. So the fix direction now closes by requiring the deliberate break: hand the store an IdentityCipher or plant a plaintext body, watch the rewritten test go RED, then restore. With the trap that makes it more than a formality, from the session that settled #344's instance 2 today: proving the INSTRUMENT can fire is only half -- the WORKLOAD must also be able to produce the failure class. Their 800-iteration repro loop returned 800/800 green against a live SQL Server while hunting a lock- contention bug, because running the two tests in isolation was the one configuration that could not generate contention. They had falsified the probe and not the rig, which felt like all of it. A rig that excludes the condition it hunts reports silence, and silence reads like evidence. Merged origin/main first (PR #148 / BACKLOG #346 landed): clean auto-merge, no conflict this time, verified by CONTENT and not only by count -- 271 items, #345, #346 and #347 all present, and all five of #347's late revisions still resolving in the merged file.
# Conflicts: # docs/adr/README.md
wshallwshall
added a commit
that referenced
this pull request
Aug 3, 2026
…ers were real but the pairing was not (#149) * docs(worktrees): the pre-squash merge-base trap, and why a three-dot diff hides it Rescuing work from an old or trailing commit is routine here, and it has a failure mode nothing in this document covered: `main` squash-merges, so a branch's own commits never become ancestors of `main`. Branch again from one of them and the new branch inherits a merge base from BEFORE the squash, so everything that landed in between is missing from it and the PR proposes deleting all of it. Measured 2026-08-02 while landing ADR 0158 from a commit pushed 1h37m after its own PR had squash-merged: 58 files and 5,726 deletions of divergence from main, conflicting on five. A three-dot diff showed two files, because three-dot resolves the merge base and the merge base is exactly what is stale. Records the two checks that do see it (`merge-base --is-ancestor` and a two-dot `diff --stat`), and that the fix is to MERGE main in rather than rebase, since the conflicting files are work that already landed via the squash. Also records why the obvious shortcut fails: a blob spot-check of a few files was run here and reported all five identical. That was true when measured and false twenty minutes later, because an armed PR touching exactly those five merged in between. It answers "are these equal now", not "will this merge". Placed under "Your PR won't merge" rather than beside the prune material, and deliberately does not restate the armed-auto-merge/BEHIND point already made at that section's third bullet. * docs(worktrees): why merge beats rebase when every commit rewrites one block The section already prescribed merge-over-rebase for the squash case, on the grounds that main's side is authoritative. There is a second and nastier reason, and it generalises beyond the squash trap. A rebase replays each commit against the new base, so a seam that every commit rewrites -- an item appended at the same EOF point -- re-raises the same conflict once per commit. The hazard is not tedium: a mid-stack resolution can keep an EARLIER DRAFT of the block, and that result carries no conflict markers, leaves git status clean, and passes a structural check, because an item that lost half its prose still has exactly one banner and still counts as one item. Nothing reports it. The rule it produces: a structural check tells you the block is COMPLETE, not that it is the version you MEANT. Verify by grepping for strings only the latest revision contains. Also records that `gh pr update-branch` cannot rescue that class -- it merges server-side, so a conflicting merge fails and the PR stays DIRTY. Complements the DIRTY row in the table above rather than restating it, and deliberately does not restate the UNKNOWN/async row, which already covers re-querying. found by: the cranky-lumiere session on docs/BACKLOG.md EOF appends, confirmed independently by the sandbox-codec session, which supplied the complete-vs-correct distinction. Routed here rather than edited in directly because docs/WORKTREES.md is contended and this section is on an open PR. * docs(worktrees): drop a provenance claim I could not verify The rebase paragraph closed with "Measured ... independently by two sessions". I did not measure it and could not verify the second session; I had it second-hand from the session that raised the finding, which has since retracted the "two sessions" wording as over-attributed -- it applied to a different fact in the same message. The finding itself is unchanged and stands on its own: the failure mode is reproducible from the description, and the paragraph already states the mechanism rather than resting on how many people saw it. What is removed is a CONFIDENCE claim about provenance, which is the one kind of sentence whose whole function is to tell the reader how much to trust the rest -- so it is the worst place to carry an unchecked number. The `gh pr update-branch` clause is deliberately left as-is. It states the mechanism (a server-side merge cannot complete a conflicting merge, so the PR stays DIRTY) rather than asserting a measurement, which is why it survives the retraction untouched. * docs(worktrees): mark the update-branch failure mode as inferred, not measured The clause asserted that `gh pr update-branch` fails on a conflict and leaves the PR DIRTY. That outcome had been relayed as measured by two sessions, then by one, then -- on audit -- by none: every session that reported it had only ever run the command against a BEHIND branch, never a DIRTY one. So the sentence had no observer at all. Checked what IS sourceable before rewriting rather than just softening it. `gh pr update-branch --help` documents the default as updating "with a merge commit (i.e., merging the base branch into the PR's branch)", and the REST endpoint takes no conflict resolution -- both real. GitHub does NOT document the endpoint's behaviour on conflict; the 422 it lists is generic, and the only conflict-adjacent note concerns a mismatched expected SHA. So the guidance stands on the mechanism, which is sound, and now SAYS it stands on the mechanism. A reader who wants to rely on the failure mode can see it was deduced from the documented default rather than observed, and weight it accordingly -- which is the whole point of the section it sits in. No attribution added: once the mechanism carries the claim, naming an observer would lend it authority it does not have, and there is no observer to name. * docs(worktrees): date the armed-auto-merge bullet, and name which check does the work Two corrections, both prompted by peer review of the section this PR adds. 1. The armed-auto-merge bullet was measured when `allow_update_branch` was `false` on this repo. It was set `true` later the same day, which may have falsified it. Deliberately NOT rewritten: GitHub's documentation does not connect that setting to base-move auto-update, and no back-fill has been observed by anyone -- replacing a stale-but-measured claim with a plausible-but-never-observed one is a strict downgrade. The bullet now carries its measurement date, states the new behaviour is unverified, and asserts nothing about back-fill. Rewrite it when someone records one. 2. The new subsection presented `--is-ancestor` and the two-dot diff as two co-equal checks. They are not. Once `--is-ancestor` passes, the merge base IS origin/main, so two-dot and three-dot compute the same thing and cannot disagree -- the diff is confirmation, not detection, and the trap only exists in the window where that check fails. The table is measured on this branch minutes apart, when a stale local checkout put it on the wrong side of the very trap it documents: three-dot reported 1 file / 50 insertions while two-dot reported 2 files / 52 insertions and 22 DELETIONS. After syncing, both read 1 file / 50 insertions. That is the failure mode a reader is most likely to miss, because on the branch they are most likely to test -- their own, up to date -- the diff agrees with itself. found by: the coordinator session for the dating (its own config change caused the drift, and it escalated rather than rewrote), and the cranky-lumiere session for the load-bearing-check distinction, while independently verifying #145. * docs(worktrees): correct the merge-base section — the numbers were real, the pairing was not The section's load-bearing illustration was wrong in three ways, all self-found, and all corrected here against re-measurement of the same commit pair (d6cb23b / 714432f): 1. "the diff shows the two files you added and nothing else" was a POST-merge three-dot reading placed beside a PRE-merge two-dot reading and presented as one comparison. The real three-dot on that pair is 13 files / 2,967 insertions / 19 deletions. 2. "the PR proposes deleting all of it" is false. A three-way merge keeps main's side of every file the branch never touched, so the 5,726 deletions were an artefact of the two-dot view, not a change anyone proposed. 3. "would have reverted a dozen merged PRs" overstated it. Re-measured: the merge base was 002be18, 11 squash-merged PRs behind, and merge-tree conflicted on exactly five files. The hazard is those conflicts and a bad resolution -- not reversion. The advice itself was sound and is kept: the branch really was stale, three-dot really does conceal staleness, --is-ancestor really is the load-bearing check, and merging main in while taking main's side really was the right repair. What changes is the framing -- the danger is restated as conflicts and a bad resolution, and the two questions are separated explicitly, since asking one and reading its answer as the other is the actual trap. Adds a short note on WHY the original survived review, because that is the transferable part: every published number was real, only the join between them was false, and nothing anywhere checks joins. It passed its author, a coordinator, an independent verification and a green CI run on that basis.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
ADR 0158 was written, committed and pushed on 2026-08-01 as
994bfb15, then never landed. The number was allocated in the coordination ledger, sodocs/adr/README.mdstopped at 0156 while 0158 was taken — the ledger pointed at a document that did not exist onmain.That gap had a cost. At least four sessions cited this silent-controls taxonomy as "ADR 0157", which is an unrelated HA demotion-safety document allocated to another worktree and still in flight on #139. The document that settles the citation was the one sitting unmerged.
Why it was stranded (the task brief had this wrong)
The originating branch did not lack a PR.
claude/intersession-communication-hooks-a52335merged as #133 at 2026-08-02T02:18:30Z;994bfb15is a trailing commit pushed ~1h37m after its own PR merged, so it missed the squash and nothing carried it afterwards.Scope
Two files.
git diff origin/mainis 475 insertions, 0 deletions, andorigin/mainis an ancestor of this branch — nothing merged is reverted.HANDOFF-announce-hook.mdfrom the same commit is deliberately not landed: it is session state, not project documentation; no root-levelHANDOFF-*.mdhas ever existed onmain; and it would publish local shim mechanics into a public repo. It stays on its own branch.Authorship
The ADR is another session's work and is preserved as such. This branch is cut from
994bfb15itself, so the original commit remains in history. Prose, voice and the author's ASCII-only convention are untouched apart from three factual corrections, each tagged inline in the ADR's own update convention rather than silently rewritten — main moved underneath the branch after it was written:main0fdc326e"0fdc326eis unreachable frommain(squash-merge); given as merged as851c849b(#130), matching the mapping the ADR already uses for7ebb2ffa/2a6649fbemail.py/direct.pycarry "the same bare"starttls()093db339(#132) gave both an explicit verifying context;alert_sinks.py:384 is the only remaining instance093db339changed the source sentence; the clause survives only inside its ownCORRECTED 2026-08-01block at:5270and as a quotation at:7476. Interval recorded; the rule is unchangedVerification
Passedlocally and--ciexit 0. 0158 is provably uncontested — exactly one commit in the repository ever added a 0158 ADR, and exactly one 0158 filename exists across all refs. The index row appears exactly once.Passed, with a positive control — an injected worktree slug made the scanner exit 1 at the injected line, so the green is evidence the detector could see that class rather than a fail-open. Detectors loaded:names=7, estate=13, site_prefixes=1.Passed.mainresolved five conflicts by taking main's side for all of them (docs/SESSION-DRIFT-CONTROLS.md,docs/WORKTREES.md,scripts/hooks/announce-session.ps1,tests/test_collision_gate.py,tests/test_coord_overlap_signals.py) — every one is coord: announce yourself to the other sessions, and stop the collision gate crying wolf #133 work already onmainvia the squash, and the still-open fix(coord): three signals that could not tell an answer from a failure #140 owns the next round on them.No engine behaviour changes.